Guangyou Lu, Chiwen Qu · Machine Learning with Applications 2026 · 2026
DOI: 10.1016/j.mlwa.2026.101026
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multimodal image-text fusion aims to integrate complementary information from visual and textual modalities to obtain more discriminative joint representations. In real-world scenarios, however, image-text classification models often suffer from unstable performance when the inputs are affected by structural degradations, such as text corruption, missing modalities, and cross-modal mismatch. These degradations change the trustworthiness of each modality and the quality of cross-modal alignment under different input conditions. We refer to this phenomenon as reliability drift. To address this problem, we propose GRaFiT: a reliability-aware image-text classification framework with multi-level guided routing. GRaFiT explicitly estimates input reliability and uses it to control adaptive fusion at three levels. At the token level, it recalibrates token representations after cross-modal fusion to emphasize more reliable local evidence. At the branch level, it predicts a dynamic coefficient α to balance different enhancement branches according to the current reliability condition. At the view level, it aggregates multi-view textual representations to suppress unreliable evidence and improve prediction stability. Experimental results show that GRaFiT significantly outperforms unimodal baselines and achieves competitive performance against representative multimodal fusion baselines. It also exhibits a relatively smooth performance decay under progressively stronger perturbations. Further analysis of internal routing behavior indicates that GRaFiT can adjust its fusion strategy according to reliability changes, and experiments on the IU-Xray dataset further demonstrate its cross-domain extensibility.
No comments yet — start the discussion below.