Emilio Dolgener Cantú, Manasi Acharya, Jim Berend, Md Golam Rasul, Jackie Ma · Companion Publication of the International Conference on Multimodal Interaction (ICMI Companion) 2026 · 2026
DOI: 10.1145/3776591.3836730
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Medical diagnosis often benefits from the side-by-side analysis of different sources of information, motivating also the development of multimodal approaches in AI-based applications for healthcare. A current technique for vision-language multimodality is a two-step process of (1) image tokenization through vector quantization followed by (2) auto-regressive modeling. In this work, we explore the image tokenization process for datasets in the medical domain, which are characterized by smaller scale and narrower semantic coverage than natural image datasets. We benchmark both reconstruction quality, codebook collapse and representational capacity of VQ-VAEs across a variety of settings, surpassing state of the art in the reconstruction task and providing a stepping stone for further development of medical multimodal auto-regressive techniques.
No comments yet — start the discussion below.