Jiayi Zou, Chaofan Chen · ACM Transactions on Multimedia Computing Communications and Applications 2026 · 2026
DOI: 10.1145/3847120
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Egocentric video reasoning capability is essential for advancing the development of first-person wearable devices. However, most existing egocentric video understanding tasks are limited to single-instance text input, overlooking the model's capability to reason about context throughout the entire duration of the video. To bridge this gap, we propose a novel task named Egocentric Video Dialogue, which refers to answering questions by considering both the video content and dialogue history. To support this task, we design a new dataset, EgoDialogue, which constructs continuous dialogues based on video question answering datasets. Egocentric videos and dialogue histories present two challenges, requiring the model to understand core interactions in visual data and the dialogue history in text, respectively. To address these issues, we introduce an EgoReason framework, consisting of an interaction exploration module and a dialogue reasoning module. The former module first reorganizes patches that correspond to the same spatial position across different frames and then sorts these blocks according to their frame order. In this way, it models spatiotemporal dynamics to help explore interactions. The latter module improves the understanding of dialogues by gradually integrating the questions and answers from each dialogue round into the fusion layer. In the experiments, our method shows superior performance compared to several VideoQA models and vision language models, achieving state-of-the-art results in prediction accuracy and various machine translation metrics.
No comments yet — start the discussion below.