The more you automate, the less you see: Hidden pitfalls of autonomous AI scientists
The rapid rise of autonomous AI scientists marks a paradigm shift in scientific discovery by automating the research lifecycle. Yet their rushed development has outpaced critical oversight, leaving key workflow decisions dangerously unscrutinized. We present a much-needed systematic analysis of open-source AI scientist systems, investigating four primary pitfalls: inappropriate benchmark selection, data leakage, metric misuse, and post hoc selection bias. Through controlled experiments that isolate each pitfall, we find systematic vulnerabilities across two representative open-source systems. Crucially, we find that these flaws are largely invisible at the level of the final manuscript, suggesting that current manuscript-centric peer review paradigms are fundamentally insufficient for ensuring the integrity of automated research. We further propose mitigation strategies and demonstrate that access to full workflow artifacts (log traces and code) enables more effective auditing. Our findings suggest that journals, conferences, and researchers should move beyond manuscript-only evaluation toward process auditing the end-to-end workflow artifacts of AI scientist systems.
Authors
- Nihar B. Shah (ORCID: https://orcid.org/0000-0001-5158-9677)
- Atoosa Kasirzadeh (ORCID: https://orcid.org/0000-0002-5967-3782)
- Ziming Luo (ORCID: https://orcid.org/0009-0005-0310-677X)
Institutions
- Carnegie Mellon University (US)
Publication Details
- Journal
- Proceedings of the National Academy of Sciences
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1073/pnas.2610214123
- Primary Topic
- Scientific Computing and Data Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00