Ivan Gentile, Motaz Saad, Kianna Kazemi, Antonella Longo · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.1870.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Standard PDF-to-Markdown pipelines silently discard embedded figures, creating a representation gap in RAG that prior work has overlooked by focusing on generator choice. We quantify this gap in the ESG domain: on 274 reports, 21.7% of 48,896 images are data-bearing yet invisible to text-only retrieval. A vision–language augmentation stage (Step3-VL-10B) recovers this content, adding 116,673 numeric tokens (+16.0%) and 1.3M words (+17.8%) across all GRI topic-standard series with 0% truncation over 8,112 descriptions. LLM-as-judge evaluation achieves 99.8% KPI-domain coverage, with 77.9% of descriptions hitting two or more KPI buckets (vs. 44.9% keyword floor). A retrieval A/B on 83 reports shows +7.0/100 completeness improvement on figure-dependent queries. These results establish the representation gap as a measurable retrieval ceiling that corpus-side enrichment must first close.
No comments yet — start the discussion below.