Georgios Chatzigeorgakidis, Dimitrios Skoutas, Alkis Simitsis · International Journal of Semantic Computing 2026 · 2026
DOI: 10.1142/s1793351x26450066
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Over the past decades, query cardinality estimation on relational databases has been proven crucial, enabling the result size estimation of all sub-plans of each query. This allowed query optimizers to design efficient query plans, towards minimizing the query response time. Lately, knowledge graphs are being preferred over traditional databases in various applications, due to their ability to represent complex relationships and semantic meaning in a domain-specific context. In this work, we present a novel approach for supervised similarity-aware cardinality estimation on knowledge graphs that incorporates similarities into the model training process, thus treating them as first-class citizens. To this end, we present SCE-KG, a model that employs a novel encoding scheme for cardinality estimation, named SG split . We further extend our approach by leveraging a Large Language Model (LLM) to generate training queries that resemble realistic user workloads, and show that combining LLM-generated and random walk samples consistently improves estimation accuracy on test workloads generated by two independent LLMs. Our experimental evaluation reveals that our approach outperforms the state-of-the-art techniques in terms of prediction error,achieving 40x and 37x improvement for star and chain queries, respectively (95th percentile), whereas its online overhead accounts for less than 0.05% of the average query execution time and less than 1% of the average query planning time, showing a sub-millisecond latency in all cases.
No comments yet — start the discussion below.