Artificial Intelligence Model’s assessment of intra-sample heterogeneity of HER2 IHC in breast cancer is related to interobserver and intraobserver agreement among pathologists

Accurate assessment of HER2 status in breast cancer has been critical for guiding therapy and has become even more important with the emergence of antibody–drug conjugates, now also indicated in HER2-low tumors. However, inter- and intraobserver variability limits the reproducibility of HER2 IHC scoring among pathologists. Artificial intelligence (AI) models offer potential to standardize and improve diagnostic accuracy and bring new insights into current practices shortcomings. We conducted a study recruiting generalist and specialist pathologists from Rede D’Or centers across Brazil to assess digitized HER2 IHC whole slide images. The same images were presented for the pathologists with an interval of one month and to the AIM-HER2 (PathAI ®, Boston, MA) AI model. Intra- and interobserver agreement, as well as agreement with AI, were measured across 126 breast cancer samples. The association between sample features and agreement metrics was also analyzed using AI spatial breakdown data. Among pathologists, the median intraobserver agreement was 66.67%, and the median agreement with AI was 60.8%. Median interobserver agreement was 67.65%, with high agreement (> 85%) in 25.4% of samples. Significant positive correlations were observed among all agreement metrics. Samples with lower intra-sample heterogeneity, as determined by AI spatial breakdown scores, were associated with higher agreement levels. Our findings highlight significant variability in HER2 IHC scoring among pathologists. We found that AI-assessed intra-sample heterogeneity is correlated with a lower agreement rate among pathologists. We believe this shows a potential limitation of current scoring practices and that AI models such as AIM-HER2 can have a role in increasing reproducibility of analysis by indicating the samples that are more likely to result in diagnostic discordance between evaluators.

Authors

Institutions

Publication Details

Journal
BMC Cancer
Published
2026-10-09
DOI
https://doi.org/10.1186/s12885-026-17049-0
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Artificial Intelligence Model’s assessment of intra-sample heterogeneity of HER2 IHC in breast cancer is related to interobserver and intraobserver agreement among pathologists

Clóvis Klock, Ana Maria da Cunha Mercante, Janaína Nagel, Antonio Alexandre de Castro et al.
BMC Cancer
AI in cancer detection
article

Artificial Intelligence Model’s assessment of intra-sample heterogeneity of HER2 IHC in breast cancer is related to interobserver and intraobserver agreement among pathologists

Clóvis Klock, Ana Maria da Cunha Mercante, Janaína Nagel, Antonio Alexandre de Castro, Annelise de Almeida Verdolin, Giselle Maria Vignal, Fernando Augusto Soares, Aline Alencar Giongo, Aline Helen da Silva Camacho, Pedro Ferrari, Amanda De Souza, Ana Lúcia de Brito Rodrigues, Patricia Barreto, Luciano Neder, Geysa Monteiro, Nicolle Gaglionone, Breast Cancer Cooperative Study Group, Gerusa Tiburzio, Ângela Karlinski, Giordano Bruno Soares-Souza, Giovanna Maximiano, Lidia C. Rezende, Ariadne Nicolli, Igor Silva, Bianca Aquino Garibaldi, Fabiano Saggioro, Andrea Cruz Monnerat, Roberto Peixoto, Mariana Morini, Isabela Werneck da Cunha, Mariana Petaccia de Macedo, Caio de Carvalho Santos, Arthur Henrique Volpato
article en

Abstract

Accurate assessment of HER2 status in breast cancer has been critical for guiding therapy and has become even more important with the emergence of antibody–drug conjugates, now also indicated in HER2-low tumors. However, inter- and intraobserver variability limits the reproducibility of HER2 IHC scoring among pathologists. Artificial intelligence (AI) models offer potential to standardize and improve diagnostic accuracy and bring new insights into current practices shortcomings. We conducted a study recruiting generalist and specialist pathologists from Rede D’Or centers across Brazil to assess digitized HER2 IHC whole slide images. The same images were presented for the pathologists with an interval of one month and to the AIM-HER2 (PathAI ®, Boston, MA) AI model. Intra- and interobserver agreement, as well as agreement with AI, were measured across 126 breast cancer samples. The association between sample features and agreement metrics was also analyzed using AI spatial breakdown data. Among pathologists, the median intraobserver agreement was 66.67%, and the median agreement with AI was 60.8%. Median interobserver agreement was 67.65%, with high agreement (> 85%) in 25.4% of samples. Significant positive correlations were observed among all agreement metrics. Samples with lower intra-sample heterogeneity, as determined by AI spatial breakdown scores, were associated with higher agreement levels. Our findings highlight significant variability in HER2 IHC scoring among pathologists. We found that AI-assessed intra-sample heterogeneity is correlated with a lower agreement rate among pathologists. We believe this shows a potential limitation of current scoring practices and that AI models such as AIM-HER2 can have a role in increasing reproducibility of analysis by indicating the samples that are more likely to result in diagnostic discordance between evaluators.

BMC Cancer
D’Or Institute for Research and Education (BR)
Openalex Percentile: Top 12%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.