Savas Yildirim, Mücahit Çevik, Ayse Basar · SN Computer Science 2026 · 2026
DOI: 10.1007/s42979-026-05302-z
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study examines the effectiveness of enhanced multilingual embedding models in improving retrieval performance for Turkish medical text data. We consider two specific medical applications in Turkish language: the TUS examination, a standardized medical assessment featuring exam questions, and Clinical QA, which involves authentic patient-physician interactions. By implementing multi-stage fine-tuning protocols on domain-specialized models, we provide detailed performance assessment and explore cross-domain transfer capabilities of the trained models. Our findings indicate that domain-specialized models improve in-domain retrieval relative to generic models, and that systematic optimization through our multi-stage pipeline yields measurable gains in retrieval precision. For instance, domain-specific fine-tuning improves TUS retrieval performance from 0.69 to 0.79 in P@1 and from 0.77 to 0.85 in MRR, while Clinical QA fine-tuning with hard-negative sampling improves P@1 from 0.33 to 0.39 and MRR from 0.41 to 0.48 relative to the vanilla encoder. In addition, reranking improves P@1 from 0.788 to 0.823 in our evaluated setting, corresponding to a 4.5% relative improvement. Furthermore, we find that, in multilingual model training, domain-specific knowledge acquired in the healthcare context of one language effectively transfers and enhances performance across other languages.
No comments yet — start the discussion below.