Comparative evaluation of traditional and advanced AI models for classifying brain tumor status from MRI reports

Abstract To evaluate automated classification of English brain MRI reports into nontumor, posttreatment-tumor, and pretreatment-tumor categories under duplicate-aware internal validation, comparing transparent lexical baselines, a recurrent deep-learning baseline, and pretrained transformer encoders with comprehensive statistical and explainability diagnostics. The analytic dataset contained 820 usable reports (541 nontumor, 181 pretreatment, 98 posttreatment) after removal of one completely blank row. Only the report Description field was used as model input. Exact and near-duplicate reports were grouped before a 585/120/115 train/validation/held-out internal test split. Models included a lexical rule baseline, TF-IDF logistic regression with unigram-bigram features (and a unigram-only sensitivity specification), a bidirectional LSTM with trainable embeddings, and four transformer encoders (BERT, BioBERT, ClinicalBERT, DistilBERT). Transformer epoch count was selected on validation data. Stochastic models (BiLSTM and transformers) were repeated using seeds 0, 42, 70, 300, and 2026. Logistic regression was assessed through feature-specification sensitivity and paired comparisons. Primary outcomes included accuracy, balanced accuracy, macro-F1, macro precision-recall AUC, class-wise metrics, duplicate-cluster bootstrap intervals, paired McNemar tests with Holm correction, calibration scores (Brier score, log loss, expected calibration error), lexical challenge tests, and quantitative LIME fidelity and stability. The unigram-bigram TF-IDF logistic regression produced the highest primary-seed test accuracy (88.70%) and macro-F1 (0.869). The unigram-only sensitivity specification achieved 86.96% accuracy and macro-F1 0.852. Among transformer models, ClinicalBERT demonstrated the strongest repeated-run profile with mean accuracy of 86.09% (SD = 0.87% points), macro-F1 of 0.831 (SD = 0.014), macro PR-AUC of 0.918 (SD = 0.003), and the most favorable calibration metrics (mean Brier score 0.252, mean ECE 0.124). BioBERT achieved 86.96% accuracy at seed 42 but showed greater variability across runs (SD = 2.51% points). BiLSTM averaged 81.22% (SD = 1.58) accuracy and macro-F1 of 0.739 (SD = 0.029). Using logistic regression as the paired reference, none of the BERT-family or BiLSTM accuracy differences remained statistically significant after Holm correction; only the comparison with the simple rule baseline was significant. Posttreatment classification remained challenging across all models (recall range: 0.333–0.815). In focal BioBERT diagnostics, LIME local fidelity ranged from 0.147 to 0.676, and top-feature stability was modest (mean pairwise top-10 Jaccard overlap approximately 0.25–0.50). No single model should be declared best overall. TF-IDF logistic regression had the strongest primary-seed point estimate, whereas ClinicalBERT was the strongest and most stable transformer across repeated runs. BioBERT was retained as the focal model for calibration, lexical challenge, and LIME diagnostics for continuity with the original analysis, not because it was the overall winner. These internally validated results support cautious use for report triage, cohort identification, or registry support, but patient-independent, center-held-out, temporal, and external validation are required before clinical deployment.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-25
DOI
https://doi.org/10.1038/s41598-026-72966-1
Primary Topic
Glioma Diagnosis and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparative evaluation of traditional and advanced AI models for classifying brain tumor status from MRI reports

Faisal Iqbal, Adven Masih, Rooha Tariq, Jabar Mahmood et al.
Scientific Reports
Glioma Diagnosis and Treatment
article

Comparative evaluation of traditional and advanced AI models for classifying brain tumor status from MRI reports

Faisal Iqbal, Adven Masih, Rooha Tariq, Jabar Mahmood, Aitizaz Ali, Daniel Musafiri Balungu, Talha Asif, Hamza Shafiq
article en

Abstract

Abstract To evaluate automated classification of English brain MRI reports into nontumor, posttreatment-tumor, and pretreatment-tumor categories under duplicate-aware internal validation, comparing transparent lexical baselines, a recurrent deep-learning baseline, and pretrained transformer encoders with comprehensive statistical and explainability diagnostics. The analytic dataset contained 820 usable reports (541 nontumor, 181 pretreatment, 98 posttreatment) after removal of one completely blank row. Only the report Description field was used as model input. Exact and near-duplicate reports were grouped before a 585/120/115 train/validation/held-out internal test split. Models included a lexical rule baseline, TF-IDF logistic regression with unigram-bigram features (and a unigram-only sensitivity specification), a bidirectional LSTM with trainable embeddings, and four transformer encoders (BERT, BioBERT, ClinicalBERT, DistilBERT). Transformer epoch count was selected on validation data. Stochastic models (BiLSTM and transformers) were repeated using seeds 0, 42, 70, 300, and 2026. Logistic regression was assessed through feature-specification sensitivity and paired comparisons. Primary outcomes included accuracy, balanced accuracy, macro-F1, macro precision-recall AUC, class-wise metrics, duplicate-cluster bootstrap intervals, paired McNemar tests with Holm correction, calibration scores (Brier score, log loss, expected calibration error), lexical challenge tests, and quantitative LIME fidelity and stability. The unigram-bigram TF-IDF logistic regression produced the highest primary-seed test accuracy (88.70%) and macro-F1 (0.869). The unigram-only sensitivity specification achieved 86.96% accuracy and macro-F1 0.852. Among transformer models, ClinicalBERT demonstrated the strongest repeated-run profile with mean accuracy of 86.09% (SD = 0.87% points), macro-F1 of 0.831 (SD = 0.014), macro PR-AUC of 0.918 (SD = 0.003), and the most favorable calibration metrics (mean Brier score 0.252, mean ECE 0.124). BioBERT achieved 86.96% accuracy at seed 42 but showed greater variability across runs (SD = 2.51% points). BiLSTM averaged 81.22% (SD = 1.58) accuracy and macro-F1 of 0.739 (SD = 0.029). Using logistic regression as the paired reference, none of the BERT-family or BiLSTM accuracy differences remained statistically significant after Holm correction; only the comparison with the simple rule baseline was significant. Posttreatment classification remained challenging across all models (recall range: 0.333–0.815). In focal BioBERT diagnostics, LIME local fidelity ranged from 0.147 to 0.676, and top-feature stability was modest (mean pairwise top-10 Jaccard overlap approximately 0.25–0.50). No single model should be declared best overall. TF-IDF logistic regression had the strongest primary-seed point estimate, whereas ClinicalBERT was the strongest and most stable transformer across repeated runs. BioBERT was retained as the focal model for calibration, lexical challenge, and LIME diagnostics for continuity with the original analysis, not because it was the overall winner. These internally validated results support cautious use for report triage, cohort identification, or registry support, but patient-independent, center-held-out, temporal, and external validation are required before clinical deployment.

Scientific Reports
Ural Federal University (RU), Karachi Medical and Dental College (PK), Asia Pacific University of Technology & Innovation (MY), University of Management and Technology (PK), İstanbul Gelişim Üniversitesi (TR)
Openalex Percentile: Top 12%
Glioma Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.