A fair multi-seed re-evaluation of mutual-information attention and Gaussian feature modeling for low-resource Chinese medical triage

Abstract Automated patient triage maps a free-text chief complaint to a clinical department and is a core component of intelligent healthcare. Recent work proposes increasingly elaborate neural mechanisms for this task in low-resource settings, but their reported gains are often established with single training runs, model-specific training protocols, and no statistical testing. We revisit one such design, MIG-BERT, which augments a Chinese BERT backbone with a Mutual Information Guided Attention (MIGA) mechanism and a Robust Gaussian Feature Modeling (RGFM) module intended to supply calibrated prediction confidence. Under a unified fine-tuning protocol with five random seeds and paired significance testing, we find that neither module produces a statistically significant improvement in weighted F1 over a properly-tuned backbone: the full model scores 0.6204 ± 0.0044 versus 0.6212 ± 0.0026 for the backbone (paired t-test, p = 0.75), and this null result replicates on two additional public datasets. We further show that, once tuned fairly, RoBERTa-wwm-ext (0.6301 ± 0.0108) and ERNIE-3.0 (0.6245 ± 0.0050) match or exceed BERT-base, reversing the counter-intuitive ranking reported previously. A quantitative uncertainty audit (Expected Calibration Error, Brier score, selective risk, and out-of-distribution detection) shows that the RGFM confidence signal is no better than a trivial maximum-probability baseline, and is worse for out-of-distribution detection (AUROC 0.30 versus 0.78). We trace the previously-reported gains to three evaluation artifacts—a module-specific learning rate, single-seed reporting, and under-tuned baselines—and release a reproducible fair-comparison benchmark of eleven model configurations and four low-resource methods across four datasets, together with an uncertainty-audit protocol. Our contribution is methodological: a rigorous, reproducible evaluation of a popular class of enhancements, and a cautionary account of how easily positive results arise in this setting.

Authors

Institutions

Publication Details

Journal
Discover Computing
Published
2026-10-06
DOI
https://doi.org/10.1007/s10791-026-10559-2
Primary Topic
Machine Learning in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A fair multi-seed re-evaluation of mutual-information attention and Gaussian feature modeling for low-resource Chinese medical triage

Yuxi Zhang
Discover Computing
Machine Learning in Healthcare
article

A fair multi-seed re-evaluation of mutual-information attention and Gaussian feature modeling for low-resource Chinese medical triage

Yuxi Zhang
article en

Abstract

Abstract Automated patient triage maps a free-text chief complaint to a clinical department and is a core component of intelligent healthcare. Recent work proposes increasingly elaborate neural mechanisms for this task in low-resource settings, but their reported gains are often established with single training runs, model-specific training protocols, and no statistical testing. We revisit one such design, MIG-BERT, which augments a Chinese BERT backbone with a Mutual Information Guided Attention (MIGA) mechanism and a Robust Gaussian Feature Modeling (RGFM) module intended to supply calibrated prediction confidence. Under a unified fine-tuning protocol with five random seeds and paired significance testing, we find that neither module produces a statistically significant improvement in weighted F1 over a properly-tuned backbone: the full model scores 0.6204 ± 0.0044 versus 0.6212 ± 0.0026 for the backbone (paired t-test, p = 0.75), and this null result replicates on two additional public datasets. We further show that, once tuned fairly, RoBERTa-wwm-ext (0.6301 ± 0.0108) and ERNIE-3.0 (0.6245 ± 0.0050) match or exceed BERT-base, reversing the counter-intuitive ranking reported previously. A quantitative uncertainty audit (Expected Calibration Error, Brier score, selective risk, and out-of-distribution detection) shows that the RGFM confidence signal is no better than a trivial maximum-probability baseline, and is worse for out-of-distribution detection (AUROC 0.30 versus 0.78). We trace the previously-reported gains to three evaluation artifacts—a module-specific learning rate, single-seed reporting, and under-tuned baselines—and release a reproducible fair-comparison benchmark of eleven model configurations and four low-resource methods across four datasets, together with an uncertainty-audit protocol. Our contribution is methodological: a rigorous, reproducible evaluation of a popular class of enhancements, and a cautionary account of how easily positive results arise in this setting.

Discover ComputingVol. 29(1)
Beijing Institute of Technology (CN)
Openalex Percentile: Top 11%
Machine Learning in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A fair multi-seed re-evaluation of mutual-information attention and Gaussian feature modeling for low-resource Chinese medical triage — Yuxi Zhang · Discover Computing (2026) | TGRS Research Map | TGRS