Can Large Language Models Evaluate Grant Proposal Quality? Revisiting the Wennerås and Wold Peer Review Data

Abstract Purpose Despite the importance of peer review for grant funding decisions, academics are often reluctant to conduct it. This can lead to long delays between submission and the final decision as well as the risk of substandard reviews from busy or non-specialist scholars. At least one funder now uses Large Language Models (LLMs) to reduce the reviewing burden but the accuracy of LLMs for scoring grant proposals needs to be assessed. Design/methodology/approach This article compares scores from a range of medium sized open-weight LLMs with peer review scores for a well-researched dataset, 142 Swedish Medical Council post-doctoral fellowship applications from 1994. Findings Whilst the LLM scores correlate moderately between each other (mean Spearman correlation: 0.34), they correlated weakly but positively and mostly statistically significantly with the average expert scores (mean Spearman correlation: 0.22). The highest rank correlation between expert scores and LLMs was 0.33 for Gemma 3 27 b based on proposal titles and summaries without their main texts, which is about half (56 %) of the correlation between reviewers. Research limitations The small sample size, old funding call and heterogeneous evaluation criteria all undermine the robustness of the analysis. Practical implications Despite the ability of LLMs to score grant proposals being quantitatively weaker than that of experts, at least in this special case, they may have role in application triage or tie-breaking. Originality/value This is the first assessment of the value of LLM scores for funding proposals.

Authors

Institutions

Publication Details

Journal
Journal of Data and Information Science
Published
2026-09-18
DOI
https://doi.org/10.1515/jdis-2026-0048
Citations
1
Primary Topic
scientometrics and bibliometrics research
Type
article
Field-Weighted Citation Impact
6.22
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Can Large Language Models Evaluate Grant Proposal Quality? Revisiting the Wennerås and Wold Peer Review Data

Ulf Sandström, Mike Thelwall
1 citations
Journal of Data and Information Science
scientometrics and bibliometrics research
6.22
article

Can Large Language Models Evaluate Grant Proposal Quality? Revisiting the Wennerås and Wold Peer Review Data

Ulf Sandström, Mike Thelwall
article en
1 citations

Abstract

Abstract Purpose Despite the importance of peer review for grant funding decisions, academics are often reluctant to conduct it. This can lead to long delays between submission and the final decision as well as the risk of substandard reviews from busy or non-specialist scholars. At least one funder now uses Large Language Models (LLMs) to reduce the reviewing burden but the accuracy of LLMs for scoring grant proposals needs to be assessed. Design/methodology/approach This article compares scores from a range of medium sized open-weight LLMs with peer review scores for a well-researched dataset, 142 Swedish Medical Council post-doctoral fellowship applications from 1994. Findings Whilst the LLM scores correlate moderately between each other (mean Spearman correlation: 0.34), they correlated weakly but positively and mostly statistically significantly with the average expert scores (mean Spearman correlation: 0.22). The highest rank correlation between expert scores and LLMs was 0.33 for Gemma 3 27 b based on proposal titles and summaries without their main texts, which is about half (56 %) of the correlation between reviewers. Research limitations The small sample size, old funding call and heterogeneous evaluation criteria all undermine the robustness of the analysis. Practical implications Despite the ability of LLMs to score grant proposals being quantitatively weaker than that of experts, at least in this special case, they may have role in application triage or tie-breaking. Originality/value This is the first assessment of the value of LLM scores for funding proposals.

Journal of Data and Information Science
KTH Royal Institute of Technology (SE), University of Sheffield (GB)
Partnerships for the goals
Openalex Percentile: Top 6%
scientometrics and bibliometrics research
6.22
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Can Large Language Models Evaluate Grant Proposal Quality? Revisiting the Wennerås and Wold Peer Review Data — Ulf Sandström, Mike Thelwall · Journal of Data and Information Science (2026) | TGRS Research Map | TGRS