Artificial intelligence in gastrointestinal endoscopy: From technical performance to clinical translation
Artificial intelligence (AI) is being increasingly used in gastrointestinal endoscopy for lesion detection, optical characterization, procedural quality monitoring, multimodal assessment, clinical decision support, and documentation. The strongest clinical evidence supports computer-aided detection during colonoscopy and selected applications for upper gastrointestinal neoplasia. Randomized trials have shown that AI can increase adenoma detection and assist in real-time lesion recognition. However, the benefits observed in controlled studies are not always reproduced in routine practice. Evidence for optical diagnosis, invasion depth assessment, inflammatory bowel disease, capsule endoscopy, pancreatobiliary endoscopy, and endoscopic ultrasonography remains heterogeneous, and many systems have been evaluated, mainly using retrospective or selected datasets. However, technical performance alone does not establish clinical value. Higher sensitivity, specificity, and diagnostic accuracy may not lead to safer decisions, more efficient workflows, or better patient outcomes. False-positive alerts, automation bias, correction burden, geographic and platform dependence, limited external validation, and poor workflow integration can affect real-world performance. Vision-foundation models, multimodal systems, and large-language models extend AI beyond single-task image analysis to reusable representations, automated reporting, and well-defined clinical support. Their clinical transferability remains uncertain because of domain shifts, data leakage, unsupported generation, and variable performance across settings. This review examined the evidence from gastrointestinal endoscopy using two complementary dimensions: clinical function and stage of evaluation. The clinical function describes the role of AI in the endoscopic pathway. Stage of evaluation distinguishes among retrospective internal testing, independent external validation, prospective clinical evaluation, interventional evaluation, and real-world or post-deployment assessment. We also use reasoning-aligned AI to describe systems whose tasks, intermediate representations, evidence integration, and outputs are organized around clinically meaningful observations and decisions, without implying human-like cognition or causal reasoning. Wider adoption should depend on independent validation, prospective evaluation of complete human–AI workflows, appropriate clinical oversight, operational reliability, governance, economic value, and reproducible benefits in routine care.
Authors
- Tongrui Yang
- Wenxin Xue
- Honggang Yu
- Xueying Wang
Institutions
- Wuhan University (CN)
- Renmin Hospital of Wuhan University (CN)
Publication Details
- Journal
- EngMedicine
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1016/j.engmed.2026.100168
- Primary Topic
- Colorectal Cancer Screening and Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00