Medical Literature-Specific Artificial Intelligence: Analyzing the performance of OpenEvidence, Heidi Evidence, Consensus, and PubMed.ai in the context of Lacrimal Drainage Disorders
PURPOSE: This study aimed to report the performance of medical literature-specific artificial intelligence (MLS-AI) deep learning tools in the context of lacrimal drainage disorders. METHODS: Prompt engineering was performed using 22 previously validated questions and statements to include common and rare aspects of lacrimal drainage disorders. Prompts were presented at least twice to the latest versions of OpenEvidence, Heidi Evidence, Consensus, and PubMed.ai [Accessed May 14-26, 2026]. The prompts included general queries, such as the management of congenital nasolacrimal duct obstructions, as well as specific queries, such as the terminology of idiopathic canalicular inflammatory disease or a history of dacryocystorhinostomy surgery. The MLS-AI were also quizzed on controversial topics, such as the use of silicone intubation and mitomycin-C in dacryocystorhinostomy. The responses were assessed for overall quality, organization, and clarity, evidence-based content, updated knowledge, specific responses, speed, and factual inaccuracies. The responses were graded by the authors as correct, partially correct, and factually incorrect. RESULTS: The overall performance of OpenEvidence and Heidi was comparable in the context of lacrimal drainage disorders. However, Heidi hallucinated twice, giving OpenEvidence an edge over Heidi. Consensus came in a close third largely due to the quality of its responses. There is much to be desired from the performance of PubMed.ai. In terms of response accuracy, the MLS-AI models performed differently. OpenEvidence's responses were graded as correct in 90.9% (20/22), and partially correct in 9.1% (2/22), with no factually incorrect responses. Heidi's responses were graded as correct in 63.6% (14/22), partially correct in 31.8% (7/22), and factually incorrect in 4.6% (1/22). In comparison, the responses of Consensus were graded as correct in 54.5% (12/22), partially correct in 40.9% (9/22), and factually incorrect in 4.6% (1/22). PubMed.ai was the worst-performing MLS-AI, with 45.5% partially correct responses, 22.7% (5/22) factually incorrect answers, and 9% (2/22) with no response. CONCLUSION: OpenEvidence has a slight edge over Heidi Evidence. PubMed.ai is lagging far behind the other 3 tools. Each of the MLS-AI tools had unique advantages and could complement one another based on context and prompt type. They need to be specifically trained further to capture more indexed mainstream journal publications. All MLS-AI responses at present should have human oversight and verification before they are accepted as true.
Authors
- Bobby S. Korn (ORCID: https://orcid.org/0000-0001-9373-131X)
- Mohammad Javed Ali
Institutions
- L V Prasad Eye Institute (IN)
Publication Details
- Journal
- Ophthalmic Plastic and Reconstructive Surgery
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1097/iop.0000000000003302
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00