Comparative evaluation of traditional and advanced AI models for classifying brain tumor status from MRI reports
Abstract To evaluate automated classification of English brain MRI reports into nontumor, posttreatment-tumor, and pretreatment-tumor categories under duplicate-aware internal validation, comparing transparent lexical baselines, a recurrent deep-learning baseline, and pretrained transformer encoders with comprehensive statistical and explainability diagnostics. The analytic dataset contained 820 usable reports (541 nontumor, 181 pretreatment, 98 posttreatment) after removal of one completely blank row. Only the report Description field was used as model input. Exact and near-duplicate reports were grouped before a 585/120/115 train/validation/held-out internal test split. Models included a lexical rule baseline, TF-IDF logistic regression with unigram-bigram features (and a unigram-only sensitivity specification), a bidirectional LSTM with trainable embeddings, and four transformer encoders (BERT, BioBERT, ClinicalBERT, DistilBERT). Transformer epoch count was selected on validation data. Stochastic models (BiLSTM and transformers) were repeated using seeds 0, 42, 70, 300, and 2026. Logistic regression was assessed through feature-specification sensitivity and paired comparisons. Primary outcomes included accuracy, balanced accuracy, macro-F1, macro precision-recall AUC, class-wise metrics, duplicate-cluster bootstrap intervals, paired McNemar tests with Holm correction, calibration scores (Brier score, log loss, expected calibration error), lexical challenge tests, and quantitative LIME fidelity and stability. The unigram-bigram TF-IDF logistic regression produced the highest primary-seed test accuracy (88.70%) and macro-F1 (0.869). The unigram-only sensitivity specification achieved 86.96% accuracy and macro-F1 0.852. Among transformer models, ClinicalBERT demonstrated the strongest repeated-run profile with mean accuracy of 86.09% (SD = 0.87% points), macro-F1 of 0.831 (SD = 0.014), macro PR-AUC of 0.918 (SD = 0.003), and the most favorable calibration metrics (mean Brier score 0.252, mean ECE 0.124). BioBERT achieved 86.96% accuracy at seed 42 but showed greater variability across runs (SD = 2.51% points). BiLSTM averaged 81.22% (SD = 1.58) accuracy and macro-F1 of 0.739 (SD = 0.029). Using logistic regression as the paired reference, none of the BERT-family or BiLSTM accuracy differences remained statistically significant after Holm correction; only the comparison with the simple rule baseline was significant. Posttreatment classification remained challenging across all models (recall range: 0.333–0.815). In focal BioBERT diagnostics, LIME local fidelity ranged from 0.147 to 0.676, and top-feature stability was modest (mean pairwise top-10 Jaccard overlap approximately 0.25–0.50). No single model should be declared best overall. TF-IDF logistic regression had the strongest primary-seed point estimate, whereas ClinicalBERT was the strongest and most stable transformer across repeated runs. BioBERT was retained as the focal model for calibration, lexical challenge, and LIME diagnostics for continuity with the original analysis, not because it was the overall winner. These internally validated results support cautious use for report triage, cohort identification, or registry support, but patient-independent, center-held-out, temporal, and external validation are required before clinical deployment.
Authors
- Faisal Iqbal (ORCID: https://orcid.org/0000-0001-6411-1226)
- Adven Masih (ORCID: https://orcid.org/0000-0001-6124-9422)
- Rooha Tariq
- Jabar Mahmood (ORCID: https://orcid.org/0000-0002-2872-0734)
- Aitizaz Ali (ORCID: https://orcid.org/0000-0002-4853-5093)
- Daniel Musafiri Balungu (ORCID: https://orcid.org/0009-0001-5098-7603)
- Talha Asif
- Hamza Shafiq
Institutions
- Ural Federal University (RU)
- Karachi Medical and Dental College (PK)
- Asia Pacific University of Technology & Innovation (MY)
- University of Management and Technology (PK)
- İstanbul Gelişim Üniversitesi (TR)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1038/s41598-026-72966-1
- Primary Topic
- Glioma Diagnosis and Treatment
- Type
- article
- Field-Weighted Citation Impact
- 0.00