Hurum Maksora Tohfa, Francisco Villaescusa-Navarro · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2609.30536
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Astronomy papers carry much of their content in figures, but literature search indexes text, so there is no way to find a plot by describing what it shows. We present FOTO, a figure-level search tool over 482,750 captions from 53,839 peer-reviewed astro-ph papers. Each figure is indexed once by its title and caption, embedded with a 110M-parameter open model that runs locally, and a large language model is applied only to the retrieved shortlist, where it verifies real figures rather than generating references to ones that do not exist. Held to the harder figure-level criterion, FOTO recovers the target for 29 to 79% of queries at recall 20 depending on register, against 6 to 20% for Pathfinder and 8 to 16% for Semantic Scholar on the easier paper-level criterion, and it beats the paid embedding API it replaced. Reranking the shortlist then raises recall at 1 from 0.667 to 0.875 on detailed queries and from 0.071 to 0.571 on vague ones. Openly available models match or beat paid LLM models for verification
No comments yet — start the discussion below.