Gene targeting in amyotrophic lateral sclerosis using causality-based feature selection and machine learning.

Kyriaki Founta Dimitra Dafou Eirini Kanata Theodoros Sklaviadis Theodoros P Zanos Anastasios Gounaris Konstantinos Xanthopoulos

Mol Med

Laboratory of Pharmacology, School of Pharmacy, School of Health Sciences, Aristotle University of Thessaloniki, 54124, Thessaloniki, Greece.

Published: January 2023

Background: Amyotrophic lateral sclerosis (ALS) is a rare progressive neurodegenerative disease that affects upper and lower motor neurons. As the molecular basis of the disease is still elusive, the development of high-throughput sequencing technologies, combined with data mining techniques and machine learning methods, could provide remarkable results in identifying pathogenetic mechanisms. High dimensionality is a major problem when applying machine learning techniques in biomedical data analysis, since a huge number of features is available for a limited number of samples. The aim of this study was to develop a methodology for training interpretable machine learning models in the classification of ALS and ALS-subtypes samples, using gene expression datasets.

Methods: We performed dimensionality reduction in gene expression data using a semi-automated preprocessing systematic gene selection procedure using Statistically Equivalent Signature (SES), a causality-based feature selection algorithm, followed by Boosted Regression Trees (XGBoost) and Random Forest to train the machine learning classifiers. The SHapley Additive exPlanations (SHAP values) were used for interpretation of the machine learning classifiers. The methodology was developed and tested using two distinct publicly available ALS RNA-seq datasets. We evaluated the performance of SES as a dimensionality reduction method against: (a) Least Absolute Shrinkage and Selection Operator (LASSO), and (b) Local Outlier Factor (LOF).

Results: The proposed methodology achieved 85.18% accuracy for the classification of cerebellum or frontal cortex samples as C9orf72-related familial ALS, sporadic ALS or healthy samples. Importantly, the genes identified as the most determinative have also been reported as disease-associated in ALS literature. When tested in the evaluation dataset, the methodology achieved 88.89% accuracy for the classification of sporadic ALS motor neuron samples. When LASSO was used as feature selection method instead of SES, the accuracy of the machine learning classifiers ranged from 74.07 to 96.30%, depending on tissue assessed, while LOF underperformed significantly (77.78% accuracy for the classification of pooled cerebellum and frontal cortex samples).

Conclusions: Using SES, we addressed the challenge of high dimensionality in gene expression data analysis, and we trained accurate machine learning ALS classifiers, specific for the gene expression patterns of different disease subtypes and tissue samples, while identifying disease-associated genes.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC9872307	PMC
http://dx.doi.org/10.1186/s10020-023-00603-y	DOI Listing

Publication Analysis

Top Keywords

machine learning

gene expression

feature selection

learning classifiers

accuracy classification

amyotrophic lateral

lateral sclerosis

causality-based feature

machine

learning

Similar Publications

Advantages of introduction of Machine learning into Patient-Controlled Anesthesia in Chronic Obstructive Pulmonary Disease and Congestive Heart Failure.

Balkan Med J

January 2025

Dow University of Health Sciences, Karachi, Pakistan.

Saim Mahmood Khan Syed Ali Hassan Faiza Rubab

View Article and Find Full Text PDF

Similar Publications

Detection of Hepatitis C Virus Infection from Patient Sera in Cell Culture Using Semi-Automated Image Analysis.

Viruses

November 2024

Department of Infectious Diseases, Molecular Virology, Section Virus-Host Interactions, Heidelberg University, 69120 Heidelberg, Germany.

Noemi Schäfer Paul Rothhaar Christian Heuss Christoph Neumann-Haefelin Robert Thimme

The study of hepatitis C virus (HCV) replication in cell culture is mainly based on cloned viral isolates requiring adaptation for efficient replication in Huh7 hepatoma cells. The analysis of wild-type (WT) isolates was enabled by the expression of SEC14L2 and by inhibitors targeting deleterious host factors. Here, we aimed to optimize cell culture models to allow infection with HCV from patient sera.

View Article and Find Full Text PDF

Similar Publications

Integrating Interpretability in Machine Learning and Deep Neural Networks: A Novel Approach to Feature Importance and Outlier Detection in COVID-19 Symptomatology and Vaccine Efficacy.

Viruses

November 2024

Faculty of Medical and Health Sciences, Tel Aviv University, Tel Aviv 6997801, Israel.

Shadi Jacob Khoury Yazeed Zoabi Mickey Scheinowitz Noam Shomron

In this study, we introduce a novel approach that integrates interpretability techniques from both traditional machine learning (ML) and deep neural networks (DNN) to quantify feature importance using global and local interpretation methods. Our method bridges the gap between interpretable ML models and powerful deep learning (DL) architectures, providing comprehensive insights into the key drivers behind model predictions, especially in detecting outliers within medical data. We applied this method to analyze COVID-19 pandemic data from 2020, yielding intriguing insights.

View Article and Find Full Text PDF

Similar Publications

Correction: Almadhor et al. AI-Driven Framework for Recognition of Guava Plant Diseases through Machine Learning from DSLR Camera Sensor Based High Resolution Imagery. 2021, , 3830.

Sensors (Basel)

December 2024

Department of Information Technology, College of Computers and Information Technology, Taif University, P.O. Box 11099, Taif 21944, Saudi Arabia.

Ahmad Almadhor Hafiz Tayyab Rauf Muhammad Ikram Ullah Lali Robertas Damaševičius Bader Alouffi

[...

View Article and Find Full Text PDF

Similar Publications

Application of Machine Learning to Predict CO Emissions in Light-Duty Vehicles.

Sensors (Basel)

December 2024

Department of Computer Science, School of Computing and Engineering, University of Huddersfield, Queensgate, Huddersfield HD1 3DH, UK.

Jeffrey Udoh Joan Lu Qiang Xu

Climate change caused by greenhouse gas (GHG) emissions is an escalating global issue, with the transportation sector being a significant contributor, accounting for approximately a quarter of all energy-related GHG emissions. In the transportation sector, vehicle emissions testing is a key part of ensuring compliance with environmental regulations. The Vehicle Certification Agency (VCA) of the UK plays a pivotal role in certifying vehicles for compliance with emissions and safety standards.

View Article and Find Full Text PDF

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!