The Alignment Paradox of Medical Large Language Models in Infertility Care: Decoupling Algorithmic Improvement From Clinical Decision-Making Quality

Abstract Background Large language models (LLMs) have been proposed as decision-support tools in assisted reproductive technology (ART), but it remains unclear whether posttraining alignment strategies translate into clinically acceptable decision support. Outcome-based benchmarks may reward token-level correctness while overlooking the reasoning quality that clinicians rely on. Objective This study aimed to evaluate whether 4 mainstream alignment paradigms for medical LLMs (supervised fine-tuning [SFT], direct preference optimization [DPO], group relative policy optimization [GRPO], and in-context learning [ICL]) produce comparable algorithmic and clinical alignment (SFT vs GRPO) when used for infertility diagnosis and treatment planning. Methods This retrospective single-center study used 8201 deidentified electronic health records from West China Second University Hospital, collected between January 2020 and December 2022 (mean age 31.79, SD 4.63 y). All 4 strategies were built on a shared open-source biomedical backbone. Evaluation comprised the following 2 tiers: (1) automatic field-level metrics (accuracy, macro- F 1 , and mean absolute error [MAE]) on 5 structured decision fields (infertility type, initial diagnosis, ART strategy, controlled ovarian stimulation [COS] regimen, and gonadotropin starting dose); and (2) blinded independent expert review by 2 reproductive medicine specialists on 100 paired cases across 4 clinical dimensions (reasoning capability, diagnostic accuracy, treatment feasibility, and hallucination). Automatic field-level evaluation included all 4 strategies, whereas blinded expert review was restricted to the SFT versus GRPO contrast. Results GRPO achieved the highest average automatic performance (eg, infertility type accuracy=92.57%, COS regimen accuracy=62.36%, ART strategy accuracy=76.49%, and gonadotropin dose MAE=44.94). However, in blinded expert review, the conservative SFT baseline showed directionally higher expert ratings than GRPO on reasoning capability and treatment feasibility; diagnostic-accuracy differences were not significant. In the 3-way best-response comparison including the original physician-charted plan, the SFT baseline was selected as the best response in 51.2% (102.3/200) of cases compared with 26.2% (52.4/200) for GRPO, and 22.6% (45.3/200) for the charted plan. This result reflects preference within the standardized review format, not evidence that model-generated decisions are clinically superior to physician decision-making. Hallucination rates were 15% (GRPO) and 18.5% (SFT), indicating that higher automatic performance did not eliminate clinically unsupported content and that neither model is ready for clinical deployment. Subgroup analyses showed GRPO improved F 1 in in vitro fertilization (IVF) and preimplantation genetic testing (PGT) but decreased F 1 in intracytoplasmic sperm injection (ICSI) cases, where male-factor information was largely captured only in unstructured fields. Conclusions Outcome-based metrics alone are insufficient proxies for clinical utility in ART decision support. Algorithmic improvement and clinical alignment suggest a possible decoupling; a phenomenon we term the alignment paradox. However, because the clinical review was based on 2 reproductive medicine specialists with marginal interrater agreement, the clinical-alignment findings should be interpreted as exploratory. External multicenter validation is required before clinical deployment.

Publication Details

Journal
Journal of Medical Internet Research
Published
2026-09-21
DOI
https://doi.org/10.2196/97221
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The Alignment Paradox of Medical Large Language Models in Infertility Care: Decoupling Algorithmic Improvement From Clinical Decision-Making Quality

Journal of Medical Internet Research
Artificial Intelligence in Healthcare and Education
article

The Alignment Paradox of Medical Large Language Models in Infertility Care: Decoupling Algorithmic Improvement From Clinical Decision-Making Quality

article en

Abstract

Abstract Background Large language models (LLMs) have been proposed as decision-support tools in assisted reproductive technology (ART), but it remains unclear whether posttraining alignment strategies translate into clinically acceptable decision support. Outcome-based benchmarks may reward token-level correctness while overlooking the reasoning quality that clinicians rely on. Objective This study aimed to evaluate whether 4 mainstream alignment paradigms for medical LLMs (supervised fine-tuning [SFT], direct preference optimization [DPO], group relative policy optimization [GRPO], and in-context learning [ICL]) produce comparable algorithmic and clinical alignment (SFT vs GRPO) when used for infertility diagnosis and treatment planning. Methods This retrospective single-center study used 8201 deidentified electronic health records from West China Second University Hospital, collected between January 2020 and December 2022 (mean age 31.79, SD 4.63 y). All 4 strategies were built on a shared open-source biomedical backbone. Evaluation comprised the following 2 tiers: (1) automatic field-level metrics (accuracy, macro- F 1 , and mean absolute error [MAE]) on 5 structured decision fields (infertility type, initial diagnosis, ART strategy, controlled ovarian stimulation [COS] regimen, and gonadotropin starting dose); and (2) blinded independent expert review by 2 reproductive medicine specialists on 100 paired cases across 4 clinical dimensions (reasoning capability, diagnostic accuracy, treatment feasibility, and hallucination). Automatic field-level evaluation included all 4 strategies, whereas blinded expert review was restricted to the SFT versus GRPO contrast. Results GRPO achieved the highest average automatic performance (eg, infertility type accuracy=92.57%, COS regimen accuracy=62.36%, ART strategy accuracy=76.49%, and gonadotropin dose MAE=44.94). However, in blinded expert review, the conservative SFT baseline showed directionally higher expert ratings than GRPO on reasoning capability and treatment feasibility; diagnostic-accuracy differences were not significant. In the 3-way best-response comparison including the original physician-charted plan, the SFT baseline was selected as the best response in 51.2% (102.3/200) of cases compared with 26.2% (52.4/200) for GRPO, and 22.6% (45.3/200) for the charted plan. This result reflects preference within the standardized review format, not evidence that model-generated decisions are clinically superior to physician decision-making. Hallucination rates were 15% (GRPO) and 18.5% (SFT), indicating that higher automatic performance did not eliminate clinically unsupported content and that neither model is ready for clinical deployment. Subgroup analyses showed GRPO improved F 1 in in vitro fertilization (IVF) and preimplantation genetic testing (PGT) but decreased F 1 in intracytoplasmic sperm injection (ICSI) cases, where male-factor information was largely captured only in unstructured fields. Conclusions Outcome-based metrics alone are insufficient proxies for clinical utility in ART decision support. Algorithmic improvement and clinical alignment suggest a possible decoupling; a phenomenon we term the alignment paradox. However, because the clinical review was based on 2 reproductive medicine specialists with marginal interrater agreement, the clinical-alignment findings should be interpreted as exploratory. External multicenter validation is required before clinical deployment.

Journal of Medical Internet ResearchVol. 28
Peace, Justice and strong institutions
Openalex Percentile: Top 100%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.