Thomas Eckart, U. Kretschmer, Erik Körner, Felix Helfer, Hagen Beelitz · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23057298
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The submission investigates how large language models (LLMs) can support the formulation of complex queries in scientific information systems, using the Federated Content Search (FCS) and its LexCQL query language as an example. Access to such systems is often hindered by implicit technical knowledge requirements, limited usability, and linguistic barriers. LLMs offer a promising alternative by enabling intuitive, natural-language interaction. A key challenge is the absence of natural-language, real-world user requests, as systems like the FCS operate exclusively with formal queries. To address this gap, synthetic information needs are generated by transforming existing query logs into natural-language user queries using LLMs, guided by simplified personas. These synthetic queries are then translated back into LexCQL, allowing both stages to be evaluated without a gold standard. Manual evaluation of 280 samples each shows that most generated information needs correctly reflect the underlying technical queries and most generated LexCQL queries are predominantly correct, confirming the feasibility of the approach. The evaluation also reveals shortcomings in the underlying technical specification. These findings will be incorporated into the next revision of the specification and demonstrate how gaps in the clarity of technical documentation can be identified that are relevant both for humans and for its use in LLM-based applications.
No comments yet — start the discussion below.