
Elena Callegari, Annika Simonsen · Northern European Journal of Language Technology 2026 · 2026
DOI: 10.3384/nejlt.2000-1533.2026.6463
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This article challenges the growing trend of using traditional linguistic tests to evaluate Large Language Models (LLMs), arguing that this approach fundamentally misunderstands both the purpose of these tests and the nature of LLMs. As a result, such evaluations often yield misleading conclusions, answering no real questions and failing to provide meaningful insights. At the core of our argument is the distinction between linguistic competence and performance: traditional linguistic tests were designed to probe human linguistic competence, yet LLMs do not acquire language through the same mechanisms and do not possess competence in the human sense. Using the analogy of birds and airplanes – both achieving flight through fundamentally different means – we illustrate why similar linguistic outputs do not imply shared underlying processes. This article calls for achange in how language-based tests are used in LLM evaluation: tests that operationalize performance can provide meaningful measures of LLM behavior, whereas tests designed to probe human linguistic competence should not be used to draw analogous competence-level conclusions about LLMs.
No comments yet — start the discussion below.