Lucas Nadolskis, Galen Pogoncheff, Michael Beyeler · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2609.16430
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Foundation-model features are increasingly used to ask what information neural activity represents, often by comparing prediction gains between nested encoding models. We show that such multimodal contrasts can change sign when only the conditioning predictor is reconstructed. Using fMRI from the Natural Scenes Dataset, DINOv2 visual features, and MPNet embeddings of MS COCO captions and Localized Narratives, a caption-narrative contrast in the additional predictive contribution of vision favors narratives when one short caption is compared with a long narrative (+0.012/+0.015 in Places), but favors captions after approximate word-count matching (-0.031/-0.023). The shift occurs across every measured ROI in both subjects and is driven primarily by differences in language-only prediction. Comparable contrasts also survive removal of image-specific content-word identity in several ROIs. These results show that nested neural contrasts do not identify represented content by themselves: predictor construction is part of the experimental design, and matched controls are required for representational claims.
No comments yet — start the discussion below.