Julian Roßkothen, David Inca Pilco, Tobias Hey, Clemens Reichmann, Ralf Reussner · KITopen 2026 · 2026
DOI: 10.5445/ir/1000196452
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Model-driven engineering tools manage large, interconnected models that may contain millions of model elements, making it increasingly difficult for engineers to locate relevant information using classical search mechanisms. Retrieval-Augmented Generation is a promising approach for enabling natural language question answering for large knowledge bases, but the applicability of text-based embedding models to highly structured, graph-shaped model artifacts has not been systematically evaluated. This paper presents an empirical study comparing 13 embedding models, 7 serialization formats, and 7 data preprocessing strategies for the retrieval of model elements. In an evaluation on two PREEvision Electric/Electronic-architecture models using automatically derived questions, we find that text embedding models are capable of capturing the semantics of model elements. However, they often only find anchor points in the model and struggle to retrieve all relevant model elements. The choice of embedding model is the dominant factor, followed by data preprocessing. Serialization has a smaller but model-dependent effect; compact, semi-structured formats such as Markdown key-value pairs perform most robustly. Embedding model retrieval-benchmark rankings transfer well to structured model element retrieval.
No comments yet — start the discussion below.