Ali Köksal, Mei Chee Leong, Vicky Sintunata, Ching Ling Chin, Wee Teck FONG · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2610.01389
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images and videos. AiSearch supports interactive search refinement through user feedback to tailor results to the user's intent, and allows visual benchmarking across multiple VLMs, enabling users to select the most suitable model for their task.
No comments yet — start the discussion below.