Zeyu Gao, Kai He, Weiheng Su, Xiaobo Pang, Inês Machado, Mercedes Jimenez‐Liñan, Brian Rous, Chunbao Wang, Chengzu Li, Chengzu Li, William McGough, Shangqi Gao, Di Zhang, Tieliang Gong, Chen Li, Faisal Mahmood, Mengling Feng, Chen Li, Chen Li, Mireia Crispin‐Ortuzar · Nature Communications 2026 · 2026
DOI: 10.1038/s41467-026-76372-z
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large Vision Language Models (LVLMs) are increasingly used in computational pathology for image classification, description generation, question answering and interactive diagnostics. However, most pathology LVLMs analyse small regions of interest rather than pyramidal, gigapixel-scale whole-slide images (WSIs), limiting their use for tasks requiring whole-slide assessment across sub-regions and magnification levels. Here, we present ALPaCA (Adapting Llama for Pathology Context Analysis), a slide-level LVLM framework for WSI question answering across diverse cancer types and tissue sites. ALPaCA is trained using 35,913 WSIs with curated descriptions and 341,051 question-answer pairs from TCGA and GTEx. It combines a LongFormer vision-text adaptor with a Gaussian mixture model-based prototyping adaptor and Llama3.1. ALPaCA exceeds 90% accuracy on internal close-ended benchmarks and maintains 77–82% accuracy on independent external cohorts. Expert pathologist evaluation of open-ended responses supports its slide-level reasoning capability. Additionally, ALPaCA can be fine-tuned on organ- or disease-specific datasets, supporting specialised pathology question answering.
No comments yet — start the discussion below.