Improving crowdsourcing for AI through cognitive-inspired data engineering

Crowdsourcing offers a fast and cost-efficient approach to obtaining human-labeled datasets. However, crowdsourced datasets and the models trained on them can inherit the cognitive constraints and biases of their annotators. In a process we refer to as cognitive-inspired data engineering, we investigate whether ideas from cognitive science can be applied to mitigate the presence of cognitive constraints and cognitive biases in crowdsourced datasets and, as a result, improve the performance of models trained on these datasets. We evaluate our approach by crowdsourcing labels for medical image diagnostic tasks using two different crowdsourcing platforms across two experiments. In Experiment 1, we collect subjective probability judgments from novice annotators through Amazon Mechanical Turk and, in Experiment 2, we collect subjective probability judgments and binary classifications from skilled annotators through DiagnosUs, a crowdsourcing platform specializing in medical and scientific data annotation. In both experiments, we find that recalibrating subjective probability judgments reduces bias, yielding more accurate crowdsourced datasets and more accurate models trained on them. Our results suggest that cognitive-inspired data engineering offers a promising avenue to improve the quality of crowdsourced datasets, with consistent downstream benefits for machine learning models.

Authors

Publication Details

Journal
Behavior Research Methods
Published
2026-09-01
DOI
https://doi.org/10.3758/s13428-026-03160-4
Citations
1
Primary Topic
Mobile Crowdsensing and Crowdsourcing
Type
article
Field-Weighted Citation Impact
10.94

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Improving crowdsourcing for AI through cognitive-inspired data engineering

Jennifer S. Trueblood, Erik Duhaime, Gunnar Paul Epping, Andrew Caplin et al.
1 citations
Behavior Research Methods
Mobile Crowdsensing and Crowdsourcing
10.94
article

Improving crowdsourcing for AI through cognitive-inspired data engineering

Jennifer S. Trueblood, Erik Duhaime, Gunnar Paul Epping, Andrew Caplin, William R. Holmes, Daniel Martin
article en
1 citations

Abstract

Crowdsourcing offers a fast and cost-efficient approach to obtaining human-labeled datasets. However, crowdsourced datasets and the models trained on them can inherit the cognitive constraints and biases of their annotators. In a process we refer to as cognitive-inspired data engineering, we investigate whether ideas from cognitive science can be applied to mitigate the presence of cognitive constraints and cognitive biases in crowdsourced datasets and, as a result, improve the performance of models trained on these datasets. We evaluate our approach by crowdsourcing labels for medical image diagnostic tasks using two different crowdsourcing platforms across two experiments. In Experiment 1, we collect subjective probability judgments from novice annotators through Amazon Mechanical Turk and, in Experiment 2, we collect subjective probability judgments and binary classifications from skilled annotators through DiagnosUs, a crowdsourcing platform specializing in medical and scientific data annotation. In both experiments, we find that recalibrating subjective probability judgments reduces bias, yielding more accurate crowdsourced datasets and more accurate models trained on them. Our results suggest that cognitive-inspired data engineering offers a promising avenue to improve the quality of crowdsourced datasets, with consistent downstream benefits for machine learning models.

Behavior Research MethodsVol. 58(10)
Alfred P. Sloan Foundation
Openalex Percentile: Top 3%
Mobile Crowdsensing and Crowdsourcing
10.94
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.