An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions

Unlabelled: Large language models (LLMs) are increasingly embedded in clinical and population health workflows, including conversational agents such as health chatbots. As chatbots evolve from rule-based approaches to hybrid and LLM-enabled designs, risks and concerns about deployment readiness shift. Unlike rule-based chatbots, LLM outputs can be unpredictable, error-prone, and difficult to validate with traditional evaluation methods. Public health teams integrating customized LLMs into interventions face practical and ethical challenges related to performance variability, uncertainties about model behaviors, and inequitable performance across languages. Although existing frameworks address domains such as safety, ethics, effectiveness, engagement, and implementation, they often assume or imply-rather than operationalize-an explicit benchmark for deployment and implementation decisions. We propose an acceptance criteria framework (ACF) to determine implementation fit, defined as meeting prespecified minimum performance standards and demonstrating nonproblematic behavior under anticipated use. The ACF uses project-relevant and off-topic prompts, structured expert review, and prespecified thresholds to produce a documented decision record that can be iteratively rerun after model revisions. We demonstrate the framework through a case application in a tobacco cessation text messaging intervention, illustrating how the ACF can guide deployment decisions.

Authors

Institutions

Publication Details

Journal
Journal of Medical Internet Research
Published
2026-07-16
DOI
https://doi.org/10.2196/92356
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions

Chelsey R. Schlechter, Leandra H. Hernández, Leticia Stevens, David W. Wetter et al.
Journal of Medical Internet Research
Artificial Intelligence in Healthcare and Education
article

An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions

Chelsey R. Schlechter, Leandra H. Hernández, Leticia Stevens, David W. Wetter, Kimberly A. Kaphingst, Andy J. King, Guilherme Del Fiol, Paul A. Estabrooks, Lindsey N. Potter, Anthony Banks, Sabrina S Thompson
article en

Abstract

Unlabelled: Large language models (LLMs) are increasingly embedded in clinical and population health workflows, including conversational agents such as health chatbots. As chatbots evolve from rule-based approaches to hybrid and LLM-enabled designs, risks and concerns about deployment readiness shift. Unlike rule-based chatbots, LLM outputs can be unpredictable, error-prone, and difficult to validate with traditional evaluation methods. Public health teams integrating customized LLMs into interventions face practical and ethical challenges related to performance variability, uncertainties about model behaviors, and inequitable performance across languages. Although existing frameworks address domains such as safety, ethics, effectiveness, engagement, and implementation, they often assume or imply-rather than operationalize-an explicit benchmark for deployment and implementation decisions. We propose an acceptance criteria framework (ACF) to determine implementation fit, defined as meeting prespecified minimum performance standards and demonstrating nonproblematic behavior under anticipated use. The ACF uses project-relevant and off-topic prompts, structured expert review, and prespecified thresholds to produce a documented decision record that can be iteratively rerun after model revisions. We demonstrate the framework through a case application in a tobacco cessation text messaging intervention, illustrating how the ACF can guide deployment decisions.

Journal of Medical Internet ResearchVol. 28
University of Utah (US), Huntsman Cancer Institute (US), Utah Department of Health (US)
Good health and well-being
Openalex Percentile: Top 12%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.