FusionNeXt-XtremeNet: An ensemble deep learning approach with LLM-Aided clinical report generation for dermoscopic image classification

Accurate classification of dermoscopic images is essential for early skin cancer detection, yet individual deep learning models exhibit architectural biases that limit performance, and translating model predictions into clinically meaningful reports requires rigorous validation and safety mechanisms. We propose FusionNeXt-XtremeNet, an ensemble deep learning framework combining ConvNeXt-Tiny, EfficientNetV2-S, and Vision Transformer (ViT-B/16) with learned weighted fusion for lesion-level diagnostic classification on dermoscopic images, integrated with a large language model-aided clinical report generation module featuring hallucination control and safety assessment. We trained the model on HAM10000 (10,015 images, 7 diagnostic categories) using patient-level stratified group 5-fold cross-validation with 3 independent seeds to prevent data leakage, and externally validated it on ISIC 2019. Statistical validation included bootstrap confidence intervals (10,000 iterations), paired t-tests with Shapiro-Wilk normality checking (Wilcoxon fallback), DeLong AUC tests, McNemar’s tests, Cochran’s Q test, and Bonferroni-Holm correction, with effect sizes (Cohen’s d, η 2 ) reported alongside all p-values. Three blinded expert annotators evaluated clinical reports using BLEU, ROUGE-L, and Likert-scale assessments, with Fleiss’ κ for inter-rater reliability. FusionNeXt-XtremeNet achieved 91.3% accuracy (95% CI: [90.9%, 91.5%]), 82.4% balanced accuracy, 76.2% macro F1-score, and 97.7% macro AUC on HAM10000, significantly outperforming all individual models (p < 0.001, Cohen’s d > 2.4 for all comparisons). On ISIC 2019, the ensemble achieved 88.2% accuracy. The clinical report module attained BLEU-4 of 0.178, ROUGE-L of 0.512, and mean clinical accuracy of 4.2/5.0 from expert evaluators (Fleiss’ κ = 0.72), with a hallucination rate of 5.2%. The proposed ensemble framework with learned weighted fusion significantly improves classification over individual architectures, and the integrated LLM-aided report generation with hallucination control provides clinically relevant decision support.

Authors

Institutions

Publication Details

Journal
Hacettepe Journal of Mathematics and Statistics
Published
2026-10-05
DOI
https://doi.org/10.15672/hujms.1947150
Primary Topic
Cutaneous Melanoma Detection and Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

FusionNeXt-XtremeNet: An ensemble deep learning approach with LLM-Aided clinical report generation for dermoscopic image classification

Jitendra Vikram Tembhurne, Saroj Shambharkar
Hacettepe Journal of Mathematics and Statistics
Cutaneous Melanoma Detection and Management
article

FusionNeXt-XtremeNet: An ensemble deep learning approach with LLM-Aided clinical report generation for dermoscopic image classification

Jitendra Vikram Tembhurne, Saroj Shambharkar
article en

Abstract

Accurate classification of dermoscopic images is essential for early skin cancer detection, yet individual deep learning models exhibit architectural biases that limit performance, and translating model predictions into clinically meaningful reports requires rigorous validation and safety mechanisms. We propose FusionNeXt-XtremeNet, an ensemble deep learning framework combining ConvNeXt-Tiny, EfficientNetV2-S, and Vision Transformer (ViT-B/16) with learned weighted fusion for lesion-level diagnostic classification on dermoscopic images, integrated with a large language model-aided clinical report generation module featuring hallucination control and safety assessment. We trained the model on HAM10000 (10,015 images, 7 diagnostic categories) using patient-level stratified group 5-fold cross-validation with 3 independent seeds to prevent data leakage, and externally validated it on ISIC 2019. Statistical validation included bootstrap confidence intervals (10,000 iterations), paired t-tests with Shapiro-Wilk normality checking (Wilcoxon fallback), DeLong AUC tests, McNemar’s tests, Cochran’s Q test, and Bonferroni-Holm correction, with effect sizes (Cohen’s d, η 2 ) reported alongside all p-values. Three blinded expert annotators evaluated clinical reports using BLEU, ROUGE-L, and Likert-scale assessments, with Fleiss’ κ for inter-rater reliability. FusionNeXt-XtremeNet achieved 91.3% accuracy (95% CI: [90.9%, 91.5%]), 82.4% balanced accuracy, 76.2% macro F1-score, and 97.7% macro AUC on HAM10000, significantly outperforming all individual models (p < 0.001, Cohen’s d > 2.4 for all comparisons). On ISIC 2019, the ensemble achieved 88.2% accuracy. The clinical report module attained BLEU-4 of 0.178, ROUGE-L of 0.512, and mean clinical accuracy of 4.2/5.0 from expert evaluators (Fleiss’ κ = 0.72), with a hallucination rate of 5.2%. The proposed ensemble framework with learned weighted fusion significantly improves classification over individual architectures, and the integrated LLM-aided report generation with hallucination control provides clinically relevant decision support.

Hacettepe Journal of Mathematics and Statistics(Advanced Online Publication)
Indian Institute of Information Technology, Nagpur (IN)
Openalex Percentile: Top 15%
Cutaneous Melanoma Detection and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.