Automated Extraction of Grade, Stage, and Quality Information From Transurethral Resection of Bladder Tumor Pathology Reports Using Natural Language Processing.

JCO Clin Cancer Inform

Alexander P. Glaser, Brian J. Jordan, Jason Cohen, Anuj Desai, Joshua J. Meeks, Feinberg School of Medicine, Northwestern University; Alexander P. Glaser, Brian J. Jordan, Joshua J. Meeks, Robert H. Lurie Comprehensive Cancer Center, Northwestern University; and Philip Silberman, Clinical and Translational Sciences Institute, Northwestern University, Chicago, IL.

Published: December 2018

Purpose: Bladder cancer is initially diagnosed and staged with a transurethral resection of bladder tumor (TURBT). Patient survival is dependent on appropriate sampling of layers of the bladder, but pathology reports are dictated as free text, making large-scale data extraction for quality improvement challenging. We sought to automate extraction of stage, grade, and quality information from TURBT pathology reports using natural language processing (NLP).

Methods: Patients undergoing TURBT were retrospectively identified using the Northwestern Enterprise Data Warehouse. An NLP algorithm was then created to extract information from free-text pathology reports and was iteratively improved using a training set of manually reviewed TURBTs. NLP accuracy was then validated using another set of manually reviewed TURBTs, and reliability was calculated using Cohen's κ.

Results: Of 3,042 TURBTs identified from 2006 to 2016, 39% were classified as benign, 35% as Ta, 11% as T1, 4% as T2, and 10% as isolated carcinoma in situ. Of 500 randomly selected manually reviewed TURBTs, NLP correctly staged 88% of specimens (κ = 0.82; 95% CI, 0.78 to 0.86). Of 272 manually reviewed T1 tumors, NLP correctly categorized grade in 100% of tumors (κ = 1), correctly categorized if muscularis propria was reported by the pathologist in 98% of tumors (κ = 0.81; 95% CI, 0.62 to 0.99), and correctly categorized if muscularis propria was present or absent in the resection specimen in 82% of tumors (κ = 0.62; 95% CI, 0.55 to 0.73). Discrepancy analysis revealed pathologist notes and deeper resection specimens as frequent reasons for NLP misclassifications.

Conclusion: We developed an NLP algorithm that demonstrates a high degree of reliability in extracting stage, grade, and presence of muscularis propria from TURBT pathology reports. Future iterations can continue to improve performance, but automated extraction of oncologic information is promising in improving quality and assisting physicians in delivery of care.

Download full-text PDF

Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC7010439PMC
http://dx.doi.org/10.1200/CCI.17.00128DOI Listing

Publication Analysis

Top Keywords

pathology reports
20
manually reviewed
16
reviewed turbts
12
correctly categorized
12
muscularis propria
12
automated extraction
8
transurethral resection
8
resection bladder
8
bladder tumor
8
reports natural
8

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!