Macoris Decena-Gimenez, Pepe Ojeda, José-Raúl Ruiz-Sarmiento, Javier González-Jiménez · Robotics 2026 · 2026
DOI: 10.3390/robotics15090175
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
High-level robotic tasks, such as those involving planning and interaction, demand a certain degree of scene understanding through suitable representations of the environment that enrich geometric information with object-level semantics, commonly referred to as semantic maps. Traditional techniques to build these maps are restricted to working with a set of categories that is fixed during training, which limits the information the resulting maps can contain. To overcome this limitation, this work presents an open-vocabulary semantic mapping pipeline that integrates TALOS (TAgging–LOcation–Segmentation) with the probabilistic, instance-aware Voxeland framework. TALOS employs a modular generative approach in which vision–language and large language models infer categories for the objects in the scene, an open-vocabulary grounding model localizes their instances, and a category-agnostic segmentation model produces pixel-accurate masks. These predictions are combined with depth data and fused into persistent 3D maps using Voxeland while preserving geometric, semantic, and instance uncertainty. The approach was evaluated on diverse ScanNet v2 scenes and compared with two alternative perception front-ends: the closed-vocabulary Detectron2/Mask R-CNN and the open-vocabulary discriminative model YOLOE. Across five independent runs, TALOS achieves a mean mAP@0.5 of 0.1607±0.0255, compared with 0.0834 for YOLOE and 0.0453 for Detectron2/Mask R-CNN, and exceeds both baselines in every run. Qualitative results further show a better balance between map completeness, geometric clarity, semantic coherence, and instance separation.
No comments yet — start the discussion below.