Tabeen Raoof · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23117196
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Vision–language models (VLMs) are increasingly deployed with smaller parameter scale and using edge hardware with constrained memory, but the kinds of spatial reasoning that survive those reductions are not characterized well, and research in the field show that edge deployment degrades accuracy but does not provide direct measurements. We evaluate Qwen2.5-VL at two scales (3B, 7B) and Gemini-2.5-Flash on the Visual Spatial Reasoning benchmark, partitioned into projective-spatial (viewpoint-dependent) and topological-containment (viewpoint-invariant) relations. An exploratory pilot (n=300) suggested that reducing model scale harms projective reasoning far more than containment. We pre-registered and ran a confirmatory test at n=2,000: the projective degradation replicated decisively (paired McNemar p ≈ 5.2 × 10⁻⁵) and containment showed no detectable degradation (p = 0.53). A logistic interaction test that treats each observation as independent found this pattern short of significance (χ²(1) = 2.67, p = 0.102); a sensitivity analysis that instead respects the paired, repeated-measures structure (item-clustered GEE) puts the same full sample at p = 0.017, but that estimate is inflated by including the 300 pilot items already known to carry a winner's-curse-inflated effect. On the 1,700 items never used to generate the hypothesis, the interaction remains directionally consistent but short of conventional significance (p ≈ 0.07–0.11 depending on method). We report both analyses rather than the single figure that best supports the hypothesis. A second pre-registered study used two one-sided tests to compare byte-identical weights on an NVIDIA Jetson Orin Nano 8GB against a desktop baseline; among inputs processed successfully by both systems, equivalence was declared within a ±3 percentage-point margin (difference +0.80 pp, 90% CI [+0.08, +1.52], p = 2.36 × 10⁻⁷, n = 1,996), a margin the data also satisfy at ±2 pp but not at ±1.5 pp. Treating the edge device's rare hard failures as unsuccessful outcomes leaves the conclusion essentially unchanged. The same device ran approximately 21× slower in its sustained regime, required an 8× context reduction to load the model, and hard-failed on 0.2% of inputs. Within this model, benchmark, and device configuration, the binding constraint on edge deployment is throughput and reliability, not accuracy.
No comments yet — start the discussion below.