AI-Assisted Scoring Improves Interobserver Agreement in Breast Cancer Biomarker Evaluation
Background: Immunohistochemical assessment of estrogen receptor (ER), progesterone receptor (PR), proliferation marker Ki67, and human epidermal growth factor receptor 2 (HER2) represents a central component of the diagnostic workflow in invasive breast carcinoma, directly influencing therapeutic decision-making, prognosis, and patient management. In this study, we evaluated the use of Indica Labs’ clinical image analysis platform for AI-assisted breast cancer biomarker scoring in a prospective observer study conducted at the University of Bern. The study compared biomarker assessment by novice observers and trained pathologists with and without AI assistance, with a focus on interobserver agreement, scoring consistency, and the potential role of AI-assisted workflows in diagnostic practice and training environments. Methods: Fifty breast cancer cases comprising 250 immunohistochemically stained slides for ER, PR, HER2, and Ki67 were independently assessed by three medical students and two board-certified pathologists in two sequential evaluation rounds. In the first round, biomarkers were scored using conventional digital pathology assessment without AI support (Visual Dx). Following a four-week washout period, the same cases were re-evaluated with assistance from HALO Breast IHC AI (AI-Assisted Dx). Observer agreement and scoring concordance were analyzed to compare performance between novice observers and experienced pathologists under unassisted and AI-assisted conditions. Results: AI-assisted assessment increased interobserver agreement at clinically relevant cutoffs across all four biomarkers in both novice observers and pathologists. Similarly, Fleiss’ kappa values improved for all biomarkers following AI assistance, with an average increase of 0.42 across observer groups. When all participants were analyzed collectively, agreement at clinical decision thresholds also improved, indicating greater concordance between novice observers and experienced pathologists during AI-assisted evaluation. Conclusions: HALO Breast IHC AI supported both novice observers and experienced pathologists by improving scoring concordance and interobserver agreement across clinically relevant breast cancer biomarkers. These findings suggest that AI-assisted workflows may contribute to greater standardization and reproducibility in biomarker assessment and may also provide supportive value in educational and training settings.
Authors
- Meredith Lodge (ORCID: https://orcid.org/0000-0001-5026-9591)
- Branislav Zagrapan (ORCID: https://orcid.org/0000-0002-1964-0580)
- Martin Wartenberg (ORCID: https://orcid.org/0000-0002-5378-3825)
- Wiebke Solaß (ORCID: https://orcid.org/0000-0002-6639-1935)
- Inti Zlobec (ORCID: https://orcid.org/0000-0001-6741-3000)
- Ben Marino Scattolo (ORCID: https://orcid.org/0009-0002-8654-2531)
- Peter Caie
- Clara Desteffani
- Florin Kalberer
- Donna Mulkern
Institutions
- University of Bern (CH)
- Bern University of Applied Sciences (CH)
- Sandia National Laboratories California (US)
Publication Details
- Journal
- Cancers
- Published
- 2026-09-16
- DOI
- https://doi.org/10.3390/cancers18183003
- Primary Topic
- AI in cancer detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00