AI, synthetic data, and the tensions of the generative gap

Synthetic data have become inseparable from the generative AI landscape as well as our algorithmic societies more broadly. The generation and deployment of synthetic data in society still remains replete with unresolved tensions and uncertainties, especially around the relationship between synthetic data and so-called “real” data. In this article I seek to make a conceptual contribution to the social science studies of synthetic data by proposing the notion of the generative gap. Here, synthetic data are defined and understood in relation to the generative gap—the productive and ongoing tensions that machine learning and AI practitioners have to negotiate between synthetic data and real data. Moreover, the generative gap simultaneously creates the problematics of synthetic data, which is explored via two sets of tension: (1) the gap between data points and (2) the gap data distributions. Ultimately, the gap between real and synthetic data is necessary and cannot be closed; rather, the generation of synthetic data is only possible insofar as it is an ongoing and contingent play of proximities and distances, a creative tension between the “too close” and the “too far away”.

Authors

Institutions

Publication Details

Journal
The Information Society
Published
2026-09-16
DOI
https://doi.org/10.1080/01972243.2026.2725355
Primary Topic
Ethics and Social Impacts of AI
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

AI, synthetic data, and the tensions of the generative gap

Benjamin N. Jacobsen
The Information Society
Ethics and Social Impacts of AI
article

AI, synthetic data, and the tensions of the generative gap

Benjamin N. Jacobsen
article en

Abstract

Synthetic data have become inseparable from the generative AI landscape as well as our algorithmic societies more broadly. The generation and deployment of synthetic data in society still remains replete with unresolved tensions and uncertainties, especially around the relationship between synthetic data and so-called “real” data. In this article I seek to make a conceptual contribution to the social science studies of synthetic data by proposing the notion of the generative gap. Here, synthetic data are defined and understood in relation to the generative gap—the productive and ongoing tensions that machine learning and AI practitioners have to negotiate between synthetic data and real data. Moreover, the generative gap simultaneously creates the problematics of synthetic data, which is explored via two sets of tension: (1) the gap between data points and (2) the gap data distributions. Ultimately, the gap between real and synthetic data is necessary and cannot be closed; rather, the generation of synthetic data is only possible insofar as it is an ongoing and contingent play of proximities and distances, a creative tension between the “too close” and the “too far away”.

The Information Society
University of York (GB)
HORIZON EUROPE European Research Council
Openalex Percentile: Top 7%
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.