Luca Cordioli, Emanuele Pucci · Companion Publication of the International Conference on Multimodal Interaction (ICMI Companion) 2026 · 2026
DOI: 10.1145/3776591.3833848
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
AI-based multimodal interaction is often framed as a move from graphical interfaces to conversation: users speak or type, and the system replies. While this shift can reduce barriers for some users, it can also reproduce exclusion when conversational systems assume normative language, typical social cues, or a single preferred modality of interaction. This position paper argues for a different interaction substrate for inclusive AI: generated state. Instead of treating text, speech, or gesture as the final interface, AI systems can translate diverse inputs into persistent, inspectable, manipulable, and multimodally renderable state objects. Such state can then be visualised as a dashboard, spoken as an audio summary, simplified into step-by-step instructions, edited through direct manipulation, queried through an agent, or transformed into alternative representations for users with different sensory, cognitive, linguistic, or social needs. Building on state generation for workplace coordination, we outline how a state-centric paradigm can support more inclusive multimodal interaction by separating the underlying coordination state from any single modality of access. We illustrate this idea through scenarios involving blind users, users with limited language proficiency, neurodivergent users, and mixed-ability teams. We conclude with design principles and research questions for inclusive state-generating AI systems.
No comments yet — start the discussion below.