Establishment of stratified reference intervals for serum CA125 using five algorithms based on real-world big data and evaluation of clinical performance

To address the lack of localized validation and precise stratification of reference intervals (RIs), this study used real-world health-checkup big data and five indirect algorithms to establish RIs for serum CA125, evaluated the consistency among the algorithms and their clinical applicability, and constructed a robust technical framework for the evidence-based assessment of RIs. A retrospective analysis of CA125 clinical data from the First Affiliated Hospital, Zhejiang University School of Medicine (FAHZU) was performed. The establishment set comprised a healthy health-checkup population studied between January 2024 and December 2025 ( n = 27,426), in whom CA125 RIs were established using traditional algorithms (EP28-NP, EP28-P) and modern algorithms (refineR, Kosmic, TMC). The evaluation set comprised health-checkup participants, outpatients and inpatients tested for CA125 between January and April 2026 and was used to analyze the discrepancy rate of the established RIs. The disease set comprised patients with eight CA125-related diseases diagnosed between January 2022 and December 2025; receiver operating characteristic (ROC) curve analysis was used to characterize the discriminative performance of CA125 across the disease cohorts and to determine optimal cut-off values in relation to the established RIs. CA125 RIs exhibited significant sex differences and dynamic perimenopausal variation. Taking the 97.5th percentile (P 97.5 ) derived from EP28-NP as an example, the established upper limits of the RIs fell into three groups: ≤19.2 U/mL for males (≥ 18 years), ≤ 31.9 U/mL for women of reproductive age (18–49 years), and a significantly lower ≤ 19.4 U/mL for postmenopausal women (≥ 50 years). RIs established by the modern algorithms were generally higher than those from the traditional algorithms; with the exception of seven TMC comparisons and one Kosmic comparison, the |BR| between the modern and traditional algorithms was < 0.375, indicating good inter-algorithm consistency. The discrepancy rates between the established and preset RIs were 6.74%, 7.01%, and 8.41% (median) for the outpatient, inpatient, and disease sets, respectively. Among the diseases examined, CA125 showed good discriminative performance for ovarian cancer overall (AUC 0.858; optimal cut-off > 29.5 U/mL), although discriminative power was limited for stage I–II disease (AUC 0.740). CA125 also provided clinically useful indications for lung cancer, pancreatic cancer, and endometriosis (AUC ≥ 0.711; optimal cut-off > 22.1 U/mL). Combining laboratory big data with multi-algorithm evaluation is a reliable approach for optimizing CA125 RIs. Sex- and age-stratified CA125 RIs may help refine the interpretation of results that fall within the borderline range, pending prospective validation, and can provide more individualized interpretation of laboratory results to support clinical risk assessment.

Authors

Institutions

Publication Details

Journal
BMC Cancer
Published
2026-09-25
DOI
https://doi.org/10.1186/s12885-026-17036-5
Primary Topic
Sepsis Diagnosis and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Establishment of stratified reference intervals for serum CA125 using five algorithms based on real-world big data and evaluation of clinical performance

Xinglun Qi, Lina Fan, Zhenzhen Tang, Dagan Yang
BMC Cancer
Sepsis Diagnosis and Treatment
article

Establishment of stratified reference intervals for serum CA125 using five algorithms based on real-world big data and evaluation of clinical performance

Xinglun Qi, Lina Fan, Zhenzhen Tang, Dagan Yang
article en

Abstract

To address the lack of localized validation and precise stratification of reference intervals (RIs), this study used real-world health-checkup big data and five indirect algorithms to establish RIs for serum CA125, evaluated the consistency among the algorithms and their clinical applicability, and constructed a robust technical framework for the evidence-based assessment of RIs. A retrospective analysis of CA125 clinical data from the First Affiliated Hospital, Zhejiang University School of Medicine (FAHZU) was performed. The establishment set comprised a healthy health-checkup population studied between January 2024 and December 2025 ( n = 27,426), in whom CA125 RIs were established using traditional algorithms (EP28-NP, EP28-P) and modern algorithms (refineR, Kosmic, TMC). The evaluation set comprised health-checkup participants, outpatients and inpatients tested for CA125 between January and April 2026 and was used to analyze the discrepancy rate of the established RIs. The disease set comprised patients with eight CA125-related diseases diagnosed between January 2022 and December 2025; receiver operating characteristic (ROC) curve analysis was used to characterize the discriminative performance of CA125 across the disease cohorts and to determine optimal cut-off values in relation to the established RIs. CA125 RIs exhibited significant sex differences and dynamic perimenopausal variation. Taking the 97.5th percentile (P 97.5 ) derived from EP28-NP as an example, the established upper limits of the RIs fell into three groups: ≤19.2 U/mL for males (≥ 18 years), ≤ 31.9 U/mL for women of reproductive age (18–49 years), and a significantly lower ≤ 19.4 U/mL for postmenopausal women (≥ 50 years). RIs established by the modern algorithms were generally higher than those from the traditional algorithms; with the exception of seven TMC comparisons and one Kosmic comparison, the |BR| between the modern and traditional algorithms was < 0.375, indicating good inter-algorithm consistency. The discrepancy rates between the established and preset RIs were 6.74%, 7.01%, and 8.41% (median) for the outpatient, inpatient, and disease sets, respectively. Among the diseases examined, CA125 showed good discriminative performance for ovarian cancer overall (AUC 0.858; optimal cut-off > 29.5 U/mL), although discriminative power was limited for stage I–II disease (AUC 0.740). CA125 also provided clinically useful indications for lung cancer, pancreatic cancer, and endometriosis (AUC ≥ 0.711; optimal cut-off > 22.1 U/mL). Combining laboratory big data with multi-algorithm evaluation is a reliable approach for optimizing CA125 RIs. Sex- and age-stratified CA125 RIs may help refine the interpretation of results that fall within the borderline range, pending prospective validation, and can provide more individualized interpretation of laboratory results to support clinical risk assessment.

BMC Cancer
First People's Hospital of Yuhang District (CN), Affiliated Zhongshan Hospital of Dalian University (CN), First Affiliated Hospital Zhejiang University (CN)
Partnerships for the goals
Openalex Percentile: Top 12%
Sepsis Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.