AI-Assisted Scoring Improves Interobserver Agreement in Breast Cancer Biomarker Evaluation

Background: Immunohistochemical assessment of estrogen receptor (ER), progesterone receptor (PR), proliferation marker Ki67, and human epidermal growth factor receptor 2 (HER2) represents a central component of the diagnostic workflow in invasive breast carcinoma, directly influencing therapeutic decision-making, prognosis, and patient management. In this study, we evaluated the use of Indica Labs’ clinical image analysis platform for AI-assisted breast cancer biomarker scoring in a prospective observer study conducted at the University of Bern. The study compared biomarker assessment by novice observers and trained pathologists with and without AI assistance, with a focus on interobserver agreement, scoring consistency, and the potential role of AI-assisted workflows in diagnostic practice and training environments. Methods: Fifty breast cancer cases comprising 250 immunohistochemically stained slides for ER, PR, HER2, and Ki67 were independently assessed by three medical students and two board-certified pathologists in two sequential evaluation rounds. In the first round, biomarkers were scored using conventional digital pathology assessment without AI support (Visual Dx). Following a four-week washout period, the same cases were re-evaluated with assistance from HALO Breast IHC AI (AI-Assisted Dx). Observer agreement and scoring concordance were analyzed to compare performance between novice observers and experienced pathologists under unassisted and AI-assisted conditions. Results: AI-assisted assessment increased interobserver agreement at clinically relevant cutoffs across all four biomarkers in both novice observers and pathologists. Similarly, Fleiss’ kappa values improved for all biomarkers following AI assistance, with an average increase of 0.42 across observer groups. When all participants were analyzed collectively, agreement at clinical decision thresholds also improved, indicating greater concordance between novice observers and experienced pathologists during AI-assisted evaluation. Conclusions: HALO Breast IHC AI supported both novice observers and experienced pathologists by improving scoring concordance and interobserver agreement across clinically relevant breast cancer biomarkers. These findings suggest that AI-assisted workflows may contribute to greater standardization and reproducibility in biomarker assessment and may also provide supportive value in educational and training settings.

Authors

Institutions

Publication Details

Journal
Cancers
Published
2026-09-16
DOI
https://doi.org/10.3390/cancers18183003
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

AI-Assisted Scoring Improves Interobserver Agreement in Breast Cancer Biomarker Evaluation

Meredith Lodge, Branislav Zagrapan, Martin Wartenberg, Wiebke Solaß et al.
Cancers
AI in cancer detection
article

AI-Assisted Scoring Improves Interobserver Agreement in Breast Cancer Biomarker Evaluation

Meredith Lodge, Branislav Zagrapan, Martin Wartenberg, Wiebke Solaß, Inti Zlobec, Ben Marino Scattolo, Peter Caie, Clara Desteffani, Florin Kalberer, Donna Mulkern
article en

Abstract

Background: Immunohistochemical assessment of estrogen receptor (ER), progesterone receptor (PR), proliferation marker Ki67, and human epidermal growth factor receptor 2 (HER2) represents a central component of the diagnostic workflow in invasive breast carcinoma, directly influencing therapeutic decision-making, prognosis, and patient management. In this study, we evaluated the use of Indica Labs’ clinical image analysis platform for AI-assisted breast cancer biomarker scoring in a prospective observer study conducted at the University of Bern. The study compared biomarker assessment by novice observers and trained pathologists with and without AI assistance, with a focus on interobserver agreement, scoring consistency, and the potential role of AI-assisted workflows in diagnostic practice and training environments. Methods: Fifty breast cancer cases comprising 250 immunohistochemically stained slides for ER, PR, HER2, and Ki67 were independently assessed by three medical students and two board-certified pathologists in two sequential evaluation rounds. In the first round, biomarkers were scored using conventional digital pathology assessment without AI support (Visual Dx). Following a four-week washout period, the same cases were re-evaluated with assistance from HALO Breast IHC AI (AI-Assisted Dx). Observer agreement and scoring concordance were analyzed to compare performance between novice observers and experienced pathologists under unassisted and AI-assisted conditions. Results: AI-assisted assessment increased interobserver agreement at clinically relevant cutoffs across all four biomarkers in both novice observers and pathologists. Similarly, Fleiss’ kappa values improved for all biomarkers following AI assistance, with an average increase of 0.42 across observer groups. When all participants were analyzed collectively, agreement at clinical decision thresholds also improved, indicating greater concordance between novice observers and experienced pathologists during AI-assisted evaluation. Conclusions: HALO Breast IHC AI supported both novice observers and experienced pathologists by improving scoring concordance and interobserver agreement across clinically relevant breast cancer biomarkers. These findings suggest that AI-assisted workflows may contribute to greater standardization and reproducibility in biomarker assessment and may also provide supportive value in educational and training settings.

CancersVol. 18(18)
University of Bern (CH), Bern University of Applied Sciences (CH), Sandia National Laboratories California (US)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.