How do deep-learning models generalize across populations? Cross-ethnicity generalization of COPD detection.

Silvia D Almeida Tobias Norajitra Carsten T Lüth Tassilo Wald Vivienn Weru Marco Nolden Paul F Jäger Oyunbileg von Stackelberg Claus Peter Heußel Oliver Weinheimer Jürgen Biederer Hans-Ulrich Kauczor Klaus Maier-Hein

Insights Imaging

Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany.

Published: August 2024

The study evaluates how well deep-learning models detect chronic obstructive pulmonary disease (COPD) in different ethnic groups, focusing on non-Hispanic Whites and African Americans.
Training on balanced datasets (both ethnic groups) and using self-supervised learning methods significantly improved model performance and reduced biases compared to using population-specific data.
The results underscore the need for equitable and effective AI healthcare solutions to ensure accurate COPD diagnosis across diverse populations.

Objectives: To evaluate the performance and potential biases of deep-learning models in detecting chronic obstructive pulmonary disease (COPD) on chest CT scans across different ethnic groups, specifically non-Hispanic White (NHW) and African American (AA) populations.

Materials And Methods: Inspiratory chest CT and clinical data from 7549 Genetic epidemiology of COPD individuals (mean age 62 years old, 56-69 interquartile range), including 5240 NHW and 2309 AA individuals, were retrospectively analyzed. Several factors influencing COPD binary classification performance on different ethnic populations were examined: (1) effects of training population: NHW-only, AA-only, balanced set (half NHW, half AA) and the entire set (NHW + AA all); (2) learning strategy: three supervised learning (SL) vs. three self-supervised learning (SSL) methods. Distribution shifts across ethnicity were further assessed for the top-performing methods.

Results: The learning strategy significantly influenced model performance, with SSL methods achieving higher performances compared to SL methods (p < 0.001), across all training configurations. Training on balanced datasets containing NHW and AA individuals resulted in improved model performance compared to population-specific datasets. Distribution shifts were found between ethnicities for the same health status, particularly when models were trained on nearest-neighbor contrastive SSL. Training on a balanced dataset resulted in fewer distribution shifts across ethnicity and health status, highlighting its efficacy in reducing biases.

Conclusion: Our findings demonstrate that utilizing SSL methods and training on large and balanced datasets can enhance COPD detection model performance and reduce biases across diverse ethnic populations. These findings emphasize the importance of equitable AI-driven healthcare solutions for COPD diagnosis.

Critical Relevance Statement: Self-supervised learning coupled with balanced datasets significantly improves COPD detection model performance, addressing biases across diverse ethnic populations and emphasizing the crucial role of equitable AI-driven healthcare solutions.

Key Points: Self-supervised learning methods outperform supervised learning methods, showing higher AUC values (p < 0.001). Balanced datasets with non-Hispanic White and African American individuals improve model performance. Training on diverse datasets enhances COPD detection accuracy. Ethnically diverse datasets reduce bias in COPD detection models. SimCLR models mitigate biases in COPD detection across ethnicities.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC11306482	PMC
http://dx.doi.org/10.1186/s13244-024-01781-x	DOI Listing

Publication Analysis

Top Keywords

deep-learning models

learning strategy

ssl methods

models generalize

generalize populations?

populations? cross-ethnicity

cross-ethnicity generalization

copd

generalization copd

copd detection

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!