Bingru Li, Han Wang, Nicholas Groom · Applied Corpus Linguistics 2026 · 2026
DOI: 10.1016/j.acorp.2026.100250
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
As the data available to corpus linguistics continues to scale ever upwards, researchers are facing a growing methodological bottleneck: while computational tools can easily count billions of words, the qualitative interpretation of these data remains a slow and labor-intensive human task. Large language models (LLMs) offer a promising way to automate this process, yet their integration into the field is often hindered by concerns over black-box unpredictability and a lack of replicability. This study introduces TACOMORE, a structured prompting framework designed to transform ad-hoc LLM interactions into a standardized linguistic protocol. Built upon four foundational principles ( Ta sk, Co ntext, Mo del, and Re plicability), the framework guides LLMs to move beyond generic probability prediction to anchoring their reasoning in the specific co-occurrence patterns of a target corpus. We applied this framework to three core corpus tasks, i.e., the analysis of keywords, collocates, and concordances, using an open corpus of COVID-19 research abstracts. After testing three popular LLMs, we find that structured prompting using TACOMORE improves accuracy and replicability, although limitations such as hallucination persist. This research offers a critical lens into the role of LLMs in corpus linguistics, highlighting their potential not only as complementary tools but also as tools for developing new methodologies within and beyond corpus linguistics.
No comments yet — start the discussion below.