Frameworks, Methodologies, and Tools for Evaluating Large Language Models in Digital Mental Health Interventions: Protocol for a Scoping Review

Abstract Background Digital mental health interventions (DMHIs) can help close persistent gaps in access to assessment, prevention, and treatment. Recent advances in generative AI, particularly large language models (LLMs), further expand this promise by enabling natural language understanding, personalization, and empathic interaction across assessment, support, and therapeutic contexts. However, significant evaluation challenges persist, including a lack of standardized constructs and validated instruments, which limit the comparability, reproducibility, and generalizability of the findings. No systematic synthesis currently documents the frameworks, methodologies, and tools used to evaluate LLMs in DMHIs, thereby hampering the development of a comprehensive evidence base to guide future evaluation efforts. Objective This scoping review aims to systematically map and synthesize the available evidence on frameworks, methodologies, and tools used to evaluate LLMs applied to DMHIs. Specifically, it aims to identify the constructs assessed, the instruments used, and the evaluation procedures and stages addressed. Methods Following the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines, this scoping review will search 5 electronic databases (PubMed, Scopus, Web of Science, IEEE Xplore, and ACM Digital Library) from January 1, 2019, to September 15, 2025. Eligibility criteria will encompass both empirical studies and theoretical proposals evaluating LLMs embedded within DMHIs. Studies limited to risk detection or decision support systems without an intervention component will be excluded. Data extraction will capture information on conceptual frameworks, methodological designs, evaluation procedures and tools, measured constructs, and other relevant contextual information. Results The systematic search was conducted between September 1 and 15, 2025, yielding 4273 records across the 5 databases. After duplicate removal, of the 4273 records, 2980 (69.7%) remained for screening. A pilot screening exercise involving 4 independent reviewers achieved high interrater reliability (free-marginal Randolph κ=0.81), with 76% (19/25 of the pilot sample) unanimous agreement, indicating adequate calibration of selection criteria. These figures are interim process indicators rather than final review findings as title and abstract screening of the remaining records is currently underway. Conclusions This scoping review is expected to provide one of the first systematic syntheses of frameworks, methodologies, and tools used to evaluate LLMs in DMHIs. By identifying prevailing patterns and gaps, the resulting evidence map is intended to serve as a practical reference for researchers, developers, and policymakers working toward a scientifically grounded, safe, ethical, and effective deployment of LLMs in mental health interventions.

Authors

Publication Details

Journal
JMIR Research Protocols
Published
2026-09-14
DOI
https://doi.org/10.2196/91677
Primary Topic
Digital Mental Health Interventions
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Frameworks, Methodologies, and Tools for Evaluating Large Language Models in Digital Mental Health Interventions: Protocol for a Scoping Review

Vania Martínez, Daniela Lira, Álvaro Jiménez-Molina, Antonio Salinas-Layana et al.
JMIR Research Protocols
Digital Mental Health Interventions
article

Frameworks, Methodologies, and Tools for Evaluating Large Language Models in Digital Mental Health Interventions: Protocol for a Scoping Review

Vania Martínez, Daniela Lira, Álvaro Jiménez-Molina, Antonio Salinas-Layana, Félix Véliz Montoya, Alexi Venegas, Rigoberto Rojas, Mario Chandía, Nicolás Muñoz
article en

Abstract

Abstract Background Digital mental health interventions (DMHIs) can help close persistent gaps in access to assessment, prevention, and treatment. Recent advances in generative AI, particularly large language models (LLMs), further expand this promise by enabling natural language understanding, personalization, and empathic interaction across assessment, support, and therapeutic contexts. However, significant evaluation challenges persist, including a lack of standardized constructs and validated instruments, which limit the comparability, reproducibility, and generalizability of the findings. No systematic synthesis currently documents the frameworks, methodologies, and tools used to evaluate LLMs in DMHIs, thereby hampering the development of a comprehensive evidence base to guide future evaluation efforts. Objective This scoping review aims to systematically map and synthesize the available evidence on frameworks, methodologies, and tools used to evaluate LLMs applied to DMHIs. Specifically, it aims to identify the constructs assessed, the instruments used, and the evaluation procedures and stages addressed. Methods Following the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines, this scoping review will search 5 electronic databases (PubMed, Scopus, Web of Science, IEEE Xplore, and ACM Digital Library) from January 1, 2019, to September 15, 2025. Eligibility criteria will encompass both empirical studies and theoretical proposals evaluating LLMs embedded within DMHIs. Studies limited to risk detection or decision support systems without an intervention component will be excluded. Data extraction will capture information on conceptual frameworks, methodological designs, evaluation procedures and tools, measured constructs, and other relevant contextual information. Results The systematic search was conducted between September 1 and 15, 2025, yielding 4273 records across the 5 databases. After duplicate removal, of the 4273 records, 2980 (69.7%) remained for screening. A pilot screening exercise involving 4 independent reviewers achieved high interrater reliability (free-marginal Randolph κ=0.81), with 76% (19/25 of the pilot sample) unanimous agreement, indicating adequate calibration of selection criteria. These figures are interim process indicators rather than final review findings as title and abstract screening of the remaining records is currently underway. Conclusions This scoping review is expected to provide one of the first systematic syntheses of frameworks, methodologies, and tools used to evaluate LLMs in DMHIs. By identifying prevailing patterns and gaps, the resulting evidence map is intended to serve as a practical reference for researchers, developers, and policymakers working toward a scientifically grounded, safe, ethical, and effective deployment of LLMs in mental health interventions.

JMIR Research ProtocolsVol. 15
Openalex Percentile: Top 9%
Digital Mental Health Interventions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.