Haojie Dai, Xiangyi Wang, Liuyi Wang, Kai Sheng, Zongtao He, Chengju Liu, Wei Ye, Qijun Chen · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2609.21504
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
DPed-VLN: A Benchmark for Socially Compliant Vision-and-Language Navigation in Dynamic Pedestrian Environments 📖 OverviewDPed-VLN is a large-scale benchmark designed for Vision-and-Language Navigation (VLN) in human-populated, dynamic indoor environments. Built upon the Habitat 3.0 simulator, it challenges embodied agents to follow natural language instructions while reacting to moving pedestrians and adhering to social-safety constraints. Unlike static VLN benchmarks, DPed-VLN introduces ORCA-controlled humanoid pedestrians and a novel hierarchical instruction protocol to decouple ordinary goal-oriented navigation from dynamic-pedestrian-aware social navigation. 📂 Dataset ContentsThis release includes the episode configurations, instruction annotations, and expert trajectories. episodes/: Navigation episode definitions (start/goal poses, pedestrian waypoints). instructions/: Level-1 and Level-2 natural language instructions. expert_trajectories/: Socially-compliant expert action sequences for Imitation Learning. 📜 CitationIf you use the DPed-VLN dataset in your research, please cite our paper and this dataset:
No comments yet — start the discussion below.