
Nathalie de Marcellis-Warin, Cristiane Melchior, Thierry Warin · Discover Artificial Intelligence 2026 · 2026
DOI: 10.1007/s44163-026-01974-x
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This article examines how large language models (LLMs) act as a new kind of ethical authority by refusing to produce certain content. Using one model (LLaMA 3.2) and one task, paraphrasing about 100,000 tweets, we study the refusals the model issues and ask what norms they enforce and how consistently. A thematic analysis of a random sample of approximately 1000 refusals, with category frequencies estimated for the full dataset, shows that they concentrate in two norm families: misinformation, reflecting a concern for truth, and hateful or harmful speech, reflecting a concern for harm. The model applies these standards selectively, on the basis of the content rather than the user, and uniformly across very different topics and contexts. We interpret this uniform, non-adaptive enforcement, with appropriate caution, as a possible driver of value homogenization, while noting that our data measure the model’s behavior rather than its effects on society. We place these findings in a brief historical context, in which ethical authority has moved from philosophers and religious institutions to states and now to AI systems, and we weigh the benefits, a higher floor for online discourse, against the risks, the global imposition of a single, largely Western-derived standard by a small number of private firms. The study contributes a social-science reading of AI refusals as empirical evidence of the ethical boundaries these systems enforce.
No comments yet — start the discussion below.