Unsupervised Extraction of Body-Text from Clinical PDF Documents.

Adel Bensahla Jamil Zaghir Christophe Gaudet-Blavignac Christian Lovis

Stud Health Technol Inform

Division of Medical Information Sciences, Geneva University Hospitals, Switzerland.

Published: August 2024

Automatic extraction of body-text within clinical PDF documents is necessary to enhance downstream NLP tasks but remains a challenge. This study presents an unsupervised algorithm designed to extract body-text leveraging large volume of data. Using DBSCAN clustering over aggregate pages, our method extracts and organize text blocks using their content and coordinates. Evaluation results demonstrate precision scores ranging from 0.82 to 0.98, recall scores from 0.62 to 0.94, and F1-scores from 0.71 to 0.96 across various medical specialty sources. Future work includes dynamic parameter adjustments for improved accuracy and using larger datasets.

Download full-text PDF	Source
http://dx.doi.org/10.3233/SHTI240382	DOI Listing

Publication Analysis

Top Keywords

extraction body-text

body-text clinical

clinical pdf

pdf documents

unsupervised extraction

documents automatic

automatic extraction

documents enhance

enhance downstream

downstream nlp

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!

A PHP Error was encountered