Weicheng Wang, Rui Lin, Jian Hu, Cong Yu · Sensors 2026 · 2026
DOI: 10.3390/s26196251
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Instance Image Navigation (IIN), an embodied navigation task, requires an agent to autonomously explore unknown environments and navigate to the vicinity of a target object indicated by an image. The core challenges are efficient semantic exploration without prior maps and accurate target discrimination among similar objects. Existing methods are mainly divided into modular approaches and 3D Gaussian splatting methods; the former lack semantic guidance, while the latter rely on pre-built maps, unsuitable for online navigation. To address this, this paper proposes Flat-IIN, a modular framework using 2D multi-channel semantic maps and a vision-language model (VLM). It constructs multiple 2D maps and integrates Multi-View Active Semantic Exploration, Global Backtracking, and VLM-based Frontier Exploration. These collaboratively leverage the maps to accomplish IIN. Experimental results on the Habitat-Matterport 3D semantic dataset show that Flat-IIN achieves a success rate of 0.794 and success weighted by normalized inverse path length of 0.331. These results demonstrate competitive navigation performance and validate its effectiveness.
No comments yet — start the discussion below.