Use of natural language processing to extract uterine weight from pathology reports

INTRODUCTION: For patients requiring hysterectomy, uterine weight is an important factor affecting surgical complexity and surgical outcomes. However, uterine weight is not readily available for analysis in administrative databases. In the present study, we use natural language processing to extract uterine weight from narrative pathology text. METHODS: Pathology text was obtained for 5000 patients sampled ±3 days of hysterectomy from hospital administrative datasets. The gross pathology subsections that describe the submitted tissues were retained for analysis. Manual annotation was performed to create a gold standard dataset. The texts were pre-processed, tokenized (split into words and punctuation), and tagged with parts-of-speech (e.g., noun, verb, number, adverb). Numeric tokens were retained for analysis. Feature engineering (e.g., variable construction and selection) included creating subspecimen flags, three tokens before and after the numeric tokens, weight units, and the prefix "weigh." Models included classification and regression trees (CART), extreme gradient boosting (XGB), and CatBoost, trained on 75%, validated on 15%, and tested in a 10% hold-out dataset. Performance on a real-world application was assessed for all hysterectomies performed over a 4-year period. RESULTS: Trained on all numeric tokens using features determined a priori, CART had 99.1% precision and 100% recall, with eight false positives. XGB performed similarly (99.3% precision; 99.9% recall). Performance was best using CatBoost's ordered target encoding and all values of the previous and subsequent three tokens (99.8% precision; 100% recall). Applied to the test dataset, CART slightly outperformed CatBoost with two fewer false positives. Summing all weights in the test set reports (n = 474), the extracted total specimen weight was a mean 9.4 g higher than the truth, owing predominantly to duplicate weights (22/37 discrepant reports). Application in the real-world dataset (n = 57 546) identified a small number of reports with data quality issues specific to low-weight specimens that should be excluded (≤10 g) or subjected to manual adjudication (>10 to ≤15 g). CONCLUSION: Natural language processing followed by simple CART can accurately extract total uterine weight from narrative pathology reports. Further work is needed to extract specific anatomic weights.

Authors

Institutions

Publication Details

Journal
International Journal of Gynecology & Obstetrics
Published
2026-10-09
DOI
https://doi.org/10.1002/ijgo.71384
Primary Topic
Biomedical Text Mining and Ontologies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Use of natural language processing to extract uterine weight from pathology reports

Ally Murji, Steven Habbous, Alinah Surani, Erik Hellsten
International Journal of Gynecology & Obstetrics
Biomedical Text Mining and Ontologies
article

Use of natural language processing to extract uterine weight from pathology reports

Ally Murji, Steven Habbous, Alinah Surani, Erik Hellsten
article en

Abstract

INTRODUCTION: For patients requiring hysterectomy, uterine weight is an important factor affecting surgical complexity and surgical outcomes. However, uterine weight is not readily available for analysis in administrative databases. In the present study, we use natural language processing to extract uterine weight from narrative pathology text. METHODS: Pathology text was obtained for 5000 patients sampled ±3 days of hysterectomy from hospital administrative datasets. The gross pathology subsections that describe the submitted tissues were retained for analysis. Manual annotation was performed to create a gold standard dataset. The texts were pre-processed, tokenized (split into words and punctuation), and tagged with parts-of-speech (e.g., noun, verb, number, adverb). Numeric tokens were retained for analysis. Feature engineering (e.g., variable construction and selection) included creating subspecimen flags, three tokens before and after the numeric tokens, weight units, and the prefix "weigh." Models included classification and regression trees (CART), extreme gradient boosting (XGB), and CatBoost, trained on 75%, validated on 15%, and tested in a 10% hold-out dataset. Performance on a real-world application was assessed for all hysterectomies performed over a 4-year period. RESULTS: Trained on all numeric tokens using features determined a priori, CART had 99.1% precision and 100% recall, with eight false positives. XGB performed similarly (99.3% precision; 99.9% recall). Performance was best using CatBoost's ordered target encoding and all values of the previous and subsequent three tokens (99.8% precision; 100% recall). Applied to the test dataset, CART slightly outperformed CatBoost with two fewer false positives. Summing all weights in the test set reports (n = 474), the extracted total specimen weight was a mean 9.4 g higher than the truth, owing predominantly to duplicate weights (22/37 discrepant reports). Application in the real-world dataset (n = 57 546) identified a small number of reports with data quality issues specific to low-weight specimens that should be excluded (≤10 g) or subjected to manual adjudication (>10 to ≤15 g). CONCLUSION: Natural language processing followed by simple CART can accurately extract total uterine weight from narrative pathology reports. Further work is needed to extract specific anatomic weights.

International Journal of Gynecology & Obstetrics
Western University (CA), University of Toronto (CA), Trillium Health Centre (CA), Ontario Health (CA)
Openalex Percentile: Top 23%
Biomedical Text Mining and Ontologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.