Xiao Zhang, Huiyuan Lai, Qianru Meng, Johan Bos · Journal of Web Semantics 2026 · 2026
DOI: 10.1016/j.websem.2026.100886
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models demonstrate strong performance on diverse tasks, yet their ability to process ontological knowledge–formal, symbolic representations of conceptual structures–remains underexplored. To address this gap, we introduce the URL Taxonomy , a framework of ontological capabilities encompassing understanding , reasoning , and learning for analyzing how language models interact with structured knowledge. Building upon this framework, we present OntoURL , the first benchmark that jointly evaluates LLMs across ontology understanding, reasoning, and learning. OntoURL comprises 15 tasks and 36,159 benchmark instances derived from 43 ontologies across eight domains. Experiments with 20 open-source LLMs reveal significant performance variations across models, tasks, and domains. Current LLMs perform well on explicit ontology understanding and several constrained reasoning tasks, but struggle with structured ontology construction, especially property-mediated relations and constraint-level outputs. Few-shot and chain-of-thought prompting yield task-dependent gains. A closed-book human evaluation shows that LLMs are competitive with or stronger than humans on closed-form understanding, reasoning, and early-stage learning tasks, whereas humans remain better at property- and constraint-level construction. These findings establish OntoURL as a reusable benchmark for evaluating LLMs on formal knowledge representations.
No comments yet — start the discussion below.