Francesco Carli, Polina Rusina, Lun Ai, Leonie Küchenhoff, Paul Ka Po To, Ellen M. McDonagh, Sebastian Lobentanzer, Fabio Petroni, Aurélien Dugourd, David Ochoa, Julio Sáez-Rodríguez · bioRxiv (Cold Spring Harbor Laboratory) 2026 · 2026
DOI: 10.64898/2026.09.01.748513
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Language models and agents are increasingly used in biomedicine, but current benchmarks reward correct answers even when the underlying reasoning is flawed. Here we introduce Karenina, an open-source framework that turns expert knowledge into multi-dimensional evaluations of questions, conversations and autonomous agents. Illustrated in Question-Answer pairs, multi-turn conversations and autonomous data-analysis, these dimensions together moves evaluation beyond scoring, enabling trustworthy decision-making with AI in biomedicine.
No comments yet — start the discussion below.