Zhipeng Yin, Zichong Wang, Min Chen, Ian Stockwell, Xin Ning, Jun Liu, Wenbin Zhang · PLOS Digital Health 2026 · 2026
DOI: 10.1371/journal.pdig.0001634
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The application of question-answering (QA) systems in the medical domain has rapidly advanced, significantly improving patients’ access to reliable health-related information. However, current approaches face notable challenges, including the difficulty in obtaining large-scale and unbiased medical datasets, significant privacy concerns, and inefficiencies due to manual dataset annotation. To address these issues, This study introduce a novel methodology leveraging publicly accessible online health forums to systematically build an unbiased, privacy-conscious QA dataset, and it employs Topic-guided Semantic Modeling (TGSM) for automated topic identification, enabling efficient and targeted annotation of relevant patient-generated content. Subsequently, this study propose a two-stage QA pipeline based on a Retriever–Reader architecture, which is further enhanced through fine-tuning state-of-the-art transformer-based models on the constructed domain-specific QA dataset. Experimental results demonstrate that our fine-tuned BioBERT significantly outperforms existing benchmarks, offering accurate patient-derived insights and providing a replicable framework for building efficient medical QA systems.
No comments yet — start the discussion below.