Enhancing selection of alcohol consumption-associated genes by random forest.

Chenglin Lyu Roby Joehanes Tianxiao Huan Daniel Levy Yi Li Mengyao Wang Xue Liu Chunyu Liu Jiantao Ma

Br J Nutr

Nutrition Epidemiology and Data Science, Friedman School of Nutrition Science and Policy, Tufts University, Boston, MA02111, USA.

Published: June 2024

Machine learning methods have been used in identifying omics markers for a variety of phenotypes. We aimed to examine whether a supervised machine learning algorithm can improve identification of alcohol-associated transcriptomic markers. In this study, we analysed array-based, whole-blood derived expression data for 17 873 gene transcripts in 5508 Framingham Heart Study participants. By using the Boruta algorithm, a supervised random forest (RF)-based feature selection method, we selected twenty-five alcohol-associated transcripts. In a testing set (30 % of entire study participants), AUC (area under the receiver operating characteristics curve) of these twenty-five transcripts were 0·73, 0·69 and 0·66 for non-drinkers . moderate drinkers, non-drinkers . heavy drinkers and moderate drinkers . heavy drinkers, respectively. The AUC of the selected transcripts by the Boruta method were comparable to those identified using conventional linear regression models, for example, AUC of 1958 transcripts identified by conventional linear regression models (false discovery rate < 0·2) were 0·74, 0·66 and 0·65, respectively. With Bonferroni correction for the twenty-five Boruta method-selected transcripts and three CVD risk factors (i.e. at < 6·7e-4), we observed thirteen transcripts were associated with obesity, three transcripts with type 2 diabetes and one transcript with hypertension. For example, we observed that alcohol consumption was inversely associated with the expression of , , and , and and were positively associated with obesity, and was inversely associated with hypertension. In conclusion, using a supervised machine learning method, the RF-based Boruta algorithm, we identified novel alcohol-associated gene transcripts.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC11216877	PMC
http://dx.doi.org/10.1017/S0007114524000795	DOI Listing

Publication Analysis

Top Keywords

machine learning

transcripts

random forest

supervised machine

gene transcripts

study participants

boruta algorithm

moderate drinkers

heavy drinkers

identified conventional

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!

A PHP Error was encountered