Redwood Meridian College · Engineering and Technology Journal 2026 · 2026
DOI: 10.47191/etj/v11i08.22
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multimodal recommenders routinely report accuracy gains over interaction-only models, yet the field rarely quantifies how much of that gain is attributable to the visual channel and how much to the textual channel. This paper reports a cross-study empirical comparison built on 54 published result cells drawn from three peer-reviewed studies that share an identical evaluation protocol on three Amazon categories: the same 5-core splits, the same released 4096-dimensional visual and 384-dimensional textual features, the same 8:1:1 per-user partition, and the same all-ranking evaluation. No new architecture is proposed. Three findings emerge. The apparent value of multimodal content depends heavily on the interaction-only reference point: 40 of 42 multimodal cells improve on a matrix-factorisation baseline by 65.91% on average, while only 30 of 42 improve on a tuned LightGCN baseline, with a mean gain of 10.29%. A published per-modality ablation shows that one content channel captures most of the available benefit, the second channel adding 3.44% on average. Published conclusions about which channel dominates conflict on identical data. Modality contribution is best read as backbone-conditional rather than as a property of the data.
No comments yet — start the discussion below.