Reproducibility of social science research using aggregate statistics with noise infused for differential privacy

Privacy-preserving analytics such as differential privacy are designed to allow the analysis of sensitive datasets while protecting individuals’ privacy. Their deployment, however, has been controversial. Critics maintain that statistical noise injected to preserve privacy can degrade the quality and feasibility of social science research. We select a benchmark of empirical findings from 93 published social science studies involving regression analyses over aggregate statistics. We evaluate whether their findings replicate on privacy noise-infused data. Under privacy budgets typical in industry, around 91% of simulated findings still support the original claims at significance level α = 0.1 . Claims based on weaker original effect sizes are more likely to be nullified or sometimes reversed. We compare distortions caused by privacy noise to those due to measurement errors and other kinds of nonsampling errors common in social statistics, and we find that the marginal impacts of privacy protection are smaller. Moreover, discrepancies due to privacy noise are often much smaller than discrepancies observed in traditional replication and robustness studies.

Authors

Institutions

Publication Details

Journal
Proceedings of the National Academy of Sciences
Published
2026-10-09
DOI
https://doi.org/10.1073/pnas.2605550123
Primary Topic
Privacy-Preserving Technologies in Data
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Reproducibility of social science research using aggregate statistics with noise infused for differential privacy

Ryan Steed, Alessandro Acquisti, Eduardo Abraham Schnadower Mustri
Proceedings of the National Academy of Sciences
Privacy-Preserving Technologies in Data
article

Reproducibility of social science research using aggregate statistics with noise infused for differential privacy

Ryan Steed, Alessandro Acquisti, Eduardo Abraham Schnadower Mustri
article en

Abstract

Privacy-preserving analytics such as differential privacy are designed to allow the analysis of sensitive datasets while protecting individuals’ privacy. Their deployment, however, has been controversial. Critics maintain that statistical noise injected to preserve privacy can degrade the quality and feasibility of social science research. We select a benchmark of empirical findings from 93 published social science studies involving regression analyses over aggregate statistics. We evaluate whether their findings replicate on privacy noise-infused data. Under privacy budgets typical in industry, around 91% of simulated findings still support the original claims at significance level α = 0.1 . Claims based on weaker original effect sizes are more likely to be nullified or sometimes reversed. We compare distortions caused by privacy noise to those due to measurement errors and other kinds of nonsampling errors common in social statistics, and we find that the marginal impacts of privacy protection are smaller. Moreover, discrepancies due to privacy noise are often much smaller than discrepancies observed in traditional replication and robustness studies.

Proceedings of the National Academy of SciencesVol. 123(41)
Massachusetts Institute of Technology (US), Carnegie Mellon University (US)
Openalex Percentile: Top 12%
Privacy-Preserving Technologies in Data
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.