EAIRA: Establishing a methodology for evaluating LLMs as scientific research assistants

Recent advancements have positioned AI, and particularly Large Language Models (LLMs) as transformative tools for scientific research, capable of addressing complex tasks that require reasoning, problem-solving, and decision-making. Their exceptional capabilities suggest their potential as scientific research assistants, but also highlight the need for holistic, rigorous, and domain-specific evaluation to assess effectiveness in real-world scientific applications. This paper describes a multifaceted methodology for Evaluating AI models as scientific Research Assistants (EAIRA) developed at Argonne National Laboratory. This methodology incorporates four primary classes of evaluations. (1) Multiple Choice Questions to assess factual recall; (2) Open Response to evaluate advanced reasoning and problem-solving skills; (3) Lab-Style Experiments involving detailed analysis of capabilities as research assistants in controlled environments; and (4) Field-Style Experiments to capture researcher-LLM interactions at scale in a wide range of scientific domains and applications. These complementary methods enable a comprehensive analysis of LLM strengths and weaknesses with respect to their scientific knowledge, reasoning abilities, and adaptability. Recognizing the rapid pace of LLM advancements, we designed the methodology to evolve and adapt so as to ensure its continued relevance and applicability. This paper describes the methodology’s state at the end of February 2025. Although developed within a subset of scientific domains, the methodology is designed to be generalizable to a wide range of scientific domains.

Authors

Institutions

Publication Details

Journal
The International Journal of High Performance Computing Applications
Published
2026-09-08
DOI
https://doi.org/10.1177/10943420261467786
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

EAIRA: Establishing a methodology for evaluating LLMs as scientific research assistants

Sandeep Madireddy, Ángel Yanguas-Gil, Ian Foster, Nicholas Chia et al.
The International Journal of High Performance Computing Applications
Artificial Intelligence in Healthcare and Education
article

EAIRA: Establishing a methodology for evaluating LLMs as scientific research assistants

Sandeep Madireddy, Ángel Yanguas-Gil, Ian Foster, Nicholas Chia, Neil Getty, Minyang Tian, Franck Cappello, Robert Underwood, Zilinghan Li, Bogdan Nicolae, Yuan-Sen Ting, Marieme Ngom, Tuan Dung Nguyen, Chenhui Zhang, Yufeng Du, Eliu Huerta, M. Mustafa Rafique, Avinash Maurya, Nesar Ramachandra, Azton Wells, Evan Antoniuk, Bo Li, Rick Stevens, Murat Keçeli, Tanwi Mallick, Bhavya Kailkhura
article en

Abstract

Recent advancements have positioned AI, and particularly Large Language Models (LLMs) as transformative tools for scientific research, capable of addressing complex tasks that require reasoning, problem-solving, and decision-making. Their exceptional capabilities suggest their potential as scientific research assistants, but also highlight the need for holistic, rigorous, and domain-specific evaluation to assess effectiveness in real-world scientific applications. This paper describes a multifaceted methodology for Evaluating AI models as scientific Research Assistants (EAIRA) developed at Argonne National Laboratory. This methodology incorporates four primary classes of evaluations. (1) Multiple Choice Questions to assess factual recall; (2) Open Response to evaluate advanced reasoning and problem-solving skills; (3) Lab-Style Experiments involving detailed analysis of capabilities as research assistants in controlled environments; and (4) Field-Style Experiments to capture researcher-LLM interactions at scale in a wide range of scientific domains and applications. These complementary methods enable a comprehensive analysis of LLM strengths and weaknesses with respect to their scientific knowledge, reasoning abilities, and adaptability. Recognizing the rapid pace of LLM advancements, we designed the methodology to evolve and adapt so as to ensure its continued relevance and applicability. This paper describes the methodology’s state at the end of February 2025. Although developed within a subset of scientific domains, the methodology is designed to be generalizable to a wide range of scientific domains.

The International Journal of High Performance Computing Applications
Argonne National Laboratory (US), Lawrence Livermore National Laboratory (US), Rochester Institute of Technology (US), University of Illinois Urbana-Champaign (US), University of Illinois Chicago (US), University of Chicago (US), The Ohio State University (US), Massachusetts Institute of Technology (US), University of Pennsylvania (US)
Argonne National Laboratory
Peace, Justice and strong institutions
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.