Large Language Models as a Clinical Interface for Predictive Modeling: A Feasibility Study in Pediatric Appendicitis

Purpose Large language models (LLMs) are increasingly used in clinical data workflows, but most clinicians lack the time, training, or support to build predictive models. We evaluated whether a ChatGPT-assisted workflow could enable clinicians to independently construct and execute a basic predictive model using routine perioperative data, using postoperative length of stay (LOS) in pediatric appendicitis as a feasibility use case. Methods We conducted a retrospective study of 228 children undergoing laparoscopic appendectomy. Ten routinely available preoperative variables were used. ChatGPT guided preprocessing and generated code for linear regression and random forest models using naïve and structured prompting. Models were trained using an 80/20 split and evaluated using MAE, RMSE, r, and R 2 . Results Naïve prompting performed poorly, while structured prompting improved model performance (MAE 0.34, RMSE 0.85, r 0.85). On the test set, linear regression achieved MAE 1.00 and random forest 0.77. Identified variables reflected expected clinical patterns. These findings demonstrate internal consistency of the workflow rather than new predictive insight. Conclusions A structured ChatGPT workflow enabled clinicians to construct standard predictive models using routine data. These findings support feasibility of clinician-directed model development, but not clinical utility, of LLM-enabled workflows.

Authors

Institutions

Publication Details

Journal
The American Surgeon
Published
2026-10-09
DOI
https://doi.org/10.1177/00031348261480836
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Large Language Models as a Clinical Interface for Predictive Modeling: A Feasibility Study in Pediatric Appendicitis

Jesse Pittard Caron, Scott W. Bloom, Jeffrey Lu, William E. Cobb et al.
The American Surgeon
Artificial Intelligence in Healthcare and Education
article

Large Language Models as a Clinical Interface for Predictive Modeling: A Feasibility Study in Pediatric Appendicitis

Jesse Pittard Caron, Scott W. Bloom, Jeffrey Lu, William E. Cobb, Daniel Rozefort, Shelby Harris, Oscar Rodriguez, Leila Raden, Lindsey Armstrong
article en

Abstract

Purpose Large language models (LLMs) are increasingly used in clinical data workflows, but most clinicians lack the time, training, or support to build predictive models. We evaluated whether a ChatGPT-assisted workflow could enable clinicians to independently construct and execute a basic predictive model using routine perioperative data, using postoperative length of stay (LOS) in pediatric appendicitis as a feasibility use case. Methods We conducted a retrospective study of 228 children undergoing laparoscopic appendectomy. Ten routinely available preoperative variables were used. ChatGPT guided preprocessing and generated code for linear regression and random forest models using naïve and structured prompting. Models were trained using an 80/20 split and evaluated using MAE, RMSE, r, and R 2 . Results Naïve prompting performed poorly, while structured prompting improved model performance (MAE 0.34, RMSE 0.85, r 0.85). On the test set, linear regression achieved MAE 1.00 and random forest 0.77. Identified variables reflected expected clinical patterns. These findings demonstrate internal consistency of the workflow rather than new predictive insight. Conclusions A structured ChatGPT workflow enabled clinicians to construct standard predictive models using routine data. These findings support feasibility of clinician-directed model development, but not clinical utility, of LLM-enabled workflows.

The American Surgeon
AdventHealth Orlando (US), AdventHealth for Children (US)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Large Language Models as a Clinical Interface for Predictive Modeling: A Feasibility Study in Pediatric Appendicitis — Jesse Pittard Caron, Scott W. Bloom, et al. · The American Surgeon (2026) | TGRS Research Map | TGRS