XingGuang Lu, Liang Kang · Discover Applied Sciences 2026 · 2026
DOI: 10.1007/s42452-026-09566-1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Spatial reasoning in multimodal large language models (MLLMs) remains vulnerable to hallucination: a model may return a fluent answer while violating metric, geometric, or cross-task constraints. This failure mode is particularly consequential when video-based spatial estimates inform navigation, manipulation, or human oversight. We present COSMES, an inference-time reliability framework for Spatial-MLLMs that connects six complementary mechanisms in one pipeline: geometry-proxy (depth-aware) frame sampling, semantic–geometric dual-branch aggregation, spatial chain-of-thought verification, geometric self-correction, hallucination-pattern detection, and stochastic consistency-based uncertainty estimation. The central contribution is the integration of input-, reasoning-, and output-level safeguards around a frozen spatial MLLM, rather than treating hallucination as only an object-presence or decoding error. COSMES returns not only a task answer but also inspectable reliability evidence, including correction records, consistency warnings, sampling dispersion, and uncertainty intervals. These outputs can provide an operational basis for abstention, re-observation, or human verification when paired with a validated downstream policy. The present evaluation is offline and does not establish calibrated confidence, broad generalization, real-time performance, or safety in deployment.
No comments yet — start the discussion below.