An integrated pipeline for prediction of Clostridioides difficile infection.

Sci Rep

Department of Molecular and Functional Genomics, Geisinger Health System, Danville, PA, USA.

Published: October 2023

With the expansion of electronic health records(EHR)-linked genomic data comes the development of machine learning-enable models. There is a pressing need to develop robust pipelines to evaluate the performance of integrated models and minimize systemic bias. We developed a prediction model of symptomatic Clostridioides difficile infection(CDI) by integrating common EHR-based and genetic risk factors(rs2227306/IL8). Our pipeline includes (1) leveraging phenotyping algorithm to minimize temporal bias, (2) performing simulation studies to determine the predictive power in samples without genetic information, (3) propensity score matching to control for the confoundings, (4) selecting machine learning algorithms to capture complex feature interactions, (5) performing oversampling to address data imbalance, and (6) optimizing models and ensuring proper bias-variance trade-off. We evaluate the performance of prediction models of CDI when including common clinical risk factors and the benefit of incorporating genetic feature(s) into the models. We emphasize the importance of building a robust integrated pipeline to avoid systemic bias and thoroughly evaluating genetic features when integrated into the prediction models in the general population and subgroups.

Download full-text PDF

Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC10545794PMC
http://dx.doi.org/10.1038/s41598-023-41753-7DOI Listing

Publication Analysis

Top Keywords

integrated pipeline
8
evaluate performance
8
systemic bias
8
prediction models
8
genetic features
8
models
6
integrated
4
prediction
4
pipeline prediction
4
prediction clostridioides
4

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!