Benjamin N. Jacobsen · The Information Society 2026 · 2026
DOI: 10.1080/01972243.2026.2725355
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Synthetic data have become inseparable from the generative AI landscape as well as our algorithmic societies more broadly. The generation and deployment of synthetic data in society still remains replete with unresolved tensions and uncertainties, especially around the relationship between synthetic data and so-called “real” data. In this article I seek to make a conceptual contribution to the social science studies of synthetic data by proposing the notion of the generative gap. Here, synthetic data are defined and understood in relation to the generative gap—the productive and ongoing tensions that machine learning and AI practitioners have to negotiate between synthetic data and real data. Moreover, the generative gap simultaneously creates the problematics of synthetic data, which is explored via two sets of tension: (1) the gap between data points and (2) the gap data distributions. Ultimately, the gap between real and synthetic data is necessary and cannot be closed; rather, the generation of synthetic data is only possible insofar as it is an ongoing and contingent play of proximities and distances, a creative tension between the “too close” and the “too far away”.
No comments yet — start the discussion below.