AI, synthetic data, and the tensions of the generative gap
Synthetic data have become inseparable from the generative AI landscape as well as our algorithmic societies more broadly. The generation and deployment of synthetic data in society still remains replete with unresolved tensions and uncertainties, especially around the relationship between synthetic data and so-called “real” data. In this article I seek to make a conceptual contribution to the social science studies of synthetic data by proposing the notion of the generative gap. Here, synthetic data are defined and understood in relation to the generative gap—the productive and ongoing tensions that machine learning and AI practitioners have to negotiate between synthetic data and real data. Moreover, the generative gap simultaneously creates the problematics of synthetic data, which is explored via two sets of tension: (1) the gap between data points and (2) the gap data distributions. Ultimately, the gap between real and synthetic data is necessary and cannot be closed; rather, the generation of synthetic data is only possible insofar as it is an ongoing and contingent play of proximities and distances, a creative tension between the “too close” and the “too far away”.
Authors
- Benjamin N. Jacobsen (ORCID: https://orcid.org/0000-0002-6656-8892)
Institutions
- University of York (GB)
Publication Details
- Journal
- The Information Society
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1080/01972243.2026.2725355
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- HORIZON EUROPE European Research Council