Frameworks, Methodologies, and Tools for Evaluating Large Language Models in Digital Mental Health Interventions: Protocol for a Scoping Review
Abstract Background Digital mental health interventions (DMHIs) can help close persistent gaps in access to assessment, prevention, and treatment. Recent advances in generative AI, particularly large language models (LLMs), further expand this promise by enabling natural language understanding, personalization, and empathic interaction across assessment, support, and therapeutic contexts. However, significant evaluation challenges persist, including a lack of standardized constructs and validated instruments, which limit the comparability, reproducibility, and generalizability of the findings. No systematic synthesis currently documents the frameworks, methodologies, and tools used to evaluate LLMs in DMHIs, thereby hampering the development of a comprehensive evidence base to guide future evaluation efforts. Objective This scoping review aims to systematically map and synthesize the available evidence on frameworks, methodologies, and tools used to evaluate LLMs applied to DMHIs. Specifically, it aims to identify the constructs assessed, the instruments used, and the evaluation procedures and stages addressed. Methods Following the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines, this scoping review will search 5 electronic databases (PubMed, Scopus, Web of Science, IEEE Xplore, and ACM Digital Library) from January 1, 2019, to September 15, 2025. Eligibility criteria will encompass both empirical studies and theoretical proposals evaluating LLMs embedded within DMHIs. Studies limited to risk detection or decision support systems without an intervention component will be excluded. Data extraction will capture information on conceptual frameworks, methodological designs, evaluation procedures and tools, measured constructs, and other relevant contextual information. Results The systematic search was conducted between September 1 and 15, 2025, yielding 4273 records across the 5 databases. After duplicate removal, of the 4273 records, 2980 (69.7%) remained for screening. A pilot screening exercise involving 4 independent reviewers achieved high interrater reliability (free-marginal Randolph κ=0.81), with 76% (19/25 of the pilot sample) unanimous agreement, indicating adequate calibration of selection criteria. These figures are interim process indicators rather than final review findings as title and abstract screening of the remaining records is currently underway. Conclusions This scoping review is expected to provide one of the first systematic syntheses of frameworks, methodologies, and tools used to evaluate LLMs in DMHIs. By identifying prevailing patterns and gaps, the resulting evidence map is intended to serve as a practical reference for researchers, developers, and policymakers working toward a scientifically grounded, safe, ethical, and effective deployment of LLMs in mental health interventions.
Authors
- Vania Martínez (ORCID: https://orcid.org/0000-0001-5980-7122)
- Daniela Lira (ORCID: https://orcid.org/0000-0001-9417-0059)
- Álvaro Jiménez-Molina (ORCID: https://orcid.org/0000-0002-5621-9322)
- Antonio Salinas-Layana
- Félix Véliz Montoya (ORCID: https://orcid.org/0009-0005-1788-2649)
- Alexi Venegas (ORCID: https://orcid.org/0009-0003-5944-4407)
- Rigoberto Rojas (ORCID: https://orcid.org/0009-0007-3202-2086)
- Mario Chandía (ORCID: https://orcid.org/0009-0006-1670-8427)
- Nicolás Muñoz (ORCID: https://orcid.org/0009-0008-6718-4998)
Publication Details
- Journal
- JMIR Research Protocols
- Published
- 2026-09-14
- DOI
- https://doi.org/10.2196/91677
- Primary Topic
- Digital Mental Health Interventions
- Type
- article
- Field-Weighted Citation Impact
- 0.00