Background:  The interest in information extraction from clinical reports for secondary data use is increasing. But experience with the productive use of information extraction processes over time is scarce. A clinical data warehouse has been in use at our university hospital for several years, which also provides an information extraction of echocardiography reports developed for general use.

Objectives:  This study aims to illustrate the difficulties encountered, while using data from a preexisting information extraction process for a large clinical study. To compare the data from the preexisting process with the data obtained from a specially developed process designed to improve the quality and completeness of the study data.

Methods:  We extracted the echocardiography variables for 440 patients from the general-use information extraction of the data warehouse (678 reports). Then we developed an information extraction process for the same variables but specifically for this study, with the aim to extract as much information as possible from the text. The extracted data of both processes were compared with a newly created gold standard defined by a cardiologist with long-standing experience in heart failure.

Results:  Among 57 echocardiography variables considered relevant for the study, 50 were documented in the routine text reports and could be extracted. Twenty of the required variables were not provided by the general-use extraction process, some others were not provided correctly. The median macro F1-score (precision, recall) across the 30 variables for which values were extracted was 0.81 (0.94, 0.77). Across all 50 variables, as relevant for the study, median macro F1-score was only 0.49 (0.56, 0.46). Employing the study-specific approach considerably improved the quality and completeness of the variables, resulting in F1-scores of 0.97 (0.98, 0.96) across all variables.

Conclusion:  Data from information extractions can be used for large clinical studies. However, preexisting information extraction processes should be treated with caution, as the time and effort spent defining each variable in the information extraction process may not be clear.

Download full-text PDF

Source
http://dx.doi.org/10.1055/s-0039-3402069DOI Listing

Publication Analysis

Top Keywords

extraction process
16
data warehouse
12
extraction
10
extraction echocardiography
8
echocardiography reports
8
variables
8
data
8
extraction processes
8
reports developed
8
data preexisting
8

Similar Publications

Winery By-Products and Effects on Atherothrombotic Markers: Focus on Platelet-Activating Factor.

Front Biosci (Landmark Ed)

January 2025

Department of Nutrition and Dietetics, School of Health Sciences and Education, Harokopio University, 17676 Athens, Greece.

Platelet aggregation and inflammation play a crucial role in atherothrombosis. Wine contains micro-constituents of proper quality and quantity that exert cardioprotective actions, partly through inhibiting platelet-activating factor (PAF), a potent inflammatory and thrombotic lipid mediator. However, wine cannot be consumed extensively due to the presence of ethanol.

View Article and Find Full Text PDF

Artificial intelligence (AI), with advantages such as automatic feature extraction and high data processing capacity and being unaffected by fatigue, can accurately analyze images obtained from colonoscopy, assess the quality of bowel preparation, and reduce the subjectivity of the operating physician, which may help to achieve standardization and normalization of colonoscopy. In this study, we aimed to explore the value of using an AI-driven intestinal image recognition model to evaluate intestinal preparation before colonoscopy. In this retrospective analysis, we analyzed the clinical data of 98 patients who underwent colonoscopy in Nantong First People's Hospital from May 2023 to October 2023.

View Article and Find Full Text PDF

This study aimed to develop gastroretentive tablets based on mucoadhesive-floating systems with encapsulated gentian (, Gentianaceae) root extract to overcome the low bioavailability and short elimination half-life of gentiopicroside, a dominant bioactive compound with systemic effect. The formulation also aimed to promote the local action of the extract in the stomach. Tablets were obtained by direct compression of sodium bicarbonate (7.

View Article and Find Full Text PDF

This study explores the development of electrospun nanofibers incorporating bioactive compounds from (Ashwagandha) root extract, focusing on optimizing extraction conditions and nanofiber composition to maximize biological activity and application potential. Using the Design of Experiment (DoE) approach, optimal extraction parameters were identified as 80% methanol, 70 °C, and 60 min, yielding high levels of phenolic compounds and antioxidant activity. Methanol concentration emerged as the critical factor influencing phytochemical properties.

View Article and Find Full Text PDF

Defatting dehulled hemp seeds is a crucial step prior to protein extraction. However, conventional methods rely on flammable solvents, posing significant health, safety, and environmental concerns. Additionally, hemp protein has poor extractability, challenging functionality, and flavor limitations, restricting its broader application in foods.

View Article and Find Full Text PDF

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!