The more you automate, the less you see: Hidden pitfalls of autonomous AI scientists

The rapid rise of autonomous AI scientists marks a paradigm shift in scientific discovery by automating the research lifecycle. Yet their rushed development has outpaced critical oversight, leaving key workflow decisions dangerously unscrutinized. We present a much-needed systematic analysis of open-source AI scientist systems, investigating four primary pitfalls: inappropriate benchmark selection, data leakage, metric misuse, and post hoc selection bias. Through controlled experiments that isolate each pitfall, we find systematic vulnerabilities across two representative open-source systems. Crucially, we find that these flaws are largely invisible at the level of the final manuscript, suggesting that current manuscript-centric peer review paradigms are fundamentally insufficient for ensuring the integrity of automated research. We further propose mitigation strategies and demonstrate that access to full workflow artifacts (log traces and code) enables more effective auditing. Our findings suggest that journals, conferences, and researchers should move beyond manuscript-only evaluation toward process auditing the end-to-end workflow artifacts of AI scientist systems.

Authors

Institutions

Publication Details

Journal
Proceedings of the National Academy of Sciences
Published
2026-10-05
DOI
https://doi.org/10.1073/pnas.2610214123
Primary Topic
Scientific Computing and Data Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

The more you automate, the less you see: Hidden pitfalls of autonomous AI scientists

Nihar B. Shah, Atoosa Kasirzadeh, Ziming Luo
Proceedings of the National Academy of Sciences
Scientific Computing and Data Management
article

The more you automate, the less you see: Hidden pitfalls of autonomous AI scientists

Nihar B. Shah, Atoosa Kasirzadeh, Ziming Luo
article en

Abstract

The rapid rise of autonomous AI scientists marks a paradigm shift in scientific discovery by automating the research lifecycle. Yet their rushed development has outpaced critical oversight, leaving key workflow decisions dangerously unscrutinized. We present a much-needed systematic analysis of open-source AI scientist systems, investigating four primary pitfalls: inappropriate benchmark selection, data leakage, metric misuse, and post hoc selection bias. Through controlled experiments that isolate each pitfall, we find systematic vulnerabilities across two representative open-source systems. Crucially, we find that these flaws are largely invisible at the level of the final manuscript, suggesting that current manuscript-centric peer review paradigms are fundamentally insufficient for ensuring the integrity of automated research. We further propose mitigation strategies and demonstrate that access to full workflow artifacts (log traces and code) enables more effective auditing. Our findings suggest that journals, conferences, and researchers should move beyond manuscript-only evaluation toward process auditing the end-to-end workflow artifacts of AI scientist systems.

Proceedings of the National Academy of SciencesVol. 123(41)
Carnegie Mellon University (US)
Openalex Percentile: Top 6%
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.