Automated extraction of key entities from non-english thorax CT reports using machine learning by large context, many-shot Generative AI
Artificial Intelligence (AI)-powered Named Entity Recognition (NER), a subfield of Natural Language Processing (NLP), can automatically analyze and extract relevant information from unstructured medical texts. However, available models are English-focused, with limited options for other languages. This study aims to identify an effective NER strategy for non-English medical reports by comparing a traditional transformer-based model against Large Language Models (LLMs) using different prompting techniques. This study performed a comparative analysis of four models for extracting key entities from anonymized Turkish Thorax Computed Tomography (CT) reports. We evaluated a spaCy-transformer model, trained with 100 reports as a baseline. This was compared against three LLM-based approaches using Google’s Gemini models: a many-shot ( n = 518 examples) prompt with Gemini 1.5 Pro, and both many-shot and five-shot prompts with Gemini 2.5 Pro. The many-shot prompts utilized a 64,000-token context. Performance was evaluated on a test set of 100 reports, focusing on five entities: anatomy (ANAT), impression (IMP), observation presence (OBS-P), absence (OBS-A), and uncertainty (OBS-U). The spaCy-transformer model achieved the highest performance in exact-match evaluation with a macro-averaged F1-score of 0.75. For relaxed-match evaluation, both the Gemini 2.5 Pro many-shot model and the spaCy-transformer achieved top-tier performance with an overall accuracy of 0.97 and a Cohen’s Kappa of 0.96. Critically, the many-shot approach with Gemini 2.5 Pro (macro F1: 0.92) significantly outperformed its five-shot counterpart (macro F1: 0.89), demonstrating the benefit of providing more examples. This study reveals that while trained transformers excel at precise boundary detection (exact match), LLMs guided by a many-shot strategy demonstrate excellent performance for relaxed-match recognition, which often carries more clinical relevance. Our results provide strong evidence that a many-shot learning approach is superior to a few-shot strategy for this task. While validated on Turkish reports, this methodology presents a promising and adaptable framework for developing high-accuracy NER tools in other languages where dedicated NLP resources are scarce.
Authors
- Efe Hasdemir
- Arzu Oğuz (ORCID: https://orcid.org/0000-0001-6512-6534)
- Burak Yağdıran (ORCID: https://orcid.org/0000-0003-0825-581X)
- Zafer Akçalı (ORCID: https://orcid.org/0000-0003-2473-4431)
- Aydan Farzaliyeva (ORCID: https://orcid.org/0000-0001-8504-4798)
- Özden Altundağ (ORCID: https://orcid.org/0000-0003-0197-6622)
- Mehmet Nezir Ramazanoğlu (ORCID: https://orcid.org/0009-0007-4976-182X)
- Murat Koçak (ORCID: https://orcid.org/0000-0001-6510-3666)
- Ahmet Muhteşem Ağıldere (ORCID: https://orcid.org/0000-0003-4223-7017)
- Fatih Guven (ORCID: https://orcid.org/0009-0006-6043-0266)
Institutions
- Başkent University (TR)
Publication Details
- Journal
- BMC Medical Informatics and Decision Making
- Published
- 2026-08-27
- DOI
- https://doi.org/10.1186/s12911-026-03784-8
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00