Identification of Phase Separating Proteins With Distributed Reduced Alphabet Representations of Sequences.

Ashwin Lahorkar Hrushikesh Bhosale Aamod Sane Vigneshwar Ramakrishnan Valadi K Jayaraman

IEEE/ACM Trans Comput Biol Bioinform

Published: April 2023

Phase separation of proteins is crucial for various cellular functions like bacterial division and tumor development, making it important to understand the molecular forces behind this process.
This research utilizes existing data to create machine learning methods, specifically Support Vector Machine and Random Forest algorithms, to predict proteins that undergo phase separation by analyzing features like hydrophobicity and amino acid flexibility.
The Random Forest model trained on a well-balanced dataset achieved a high accuracy of 97%, demonstrating that using interpretable features can enhance the prediction of phase separating proteins and potentially inform disease treatment strategies.

Phase separation of proteins play key roles in cellular physiology including bacterial division, tumorigenesis etc. Consequently, understanding the molecular forces that drive phase separation has gained considerable attention and several factors including hydrophobicity, protein dynamics, etc., have been implicated in phase separation. Data-driven identification of new phase separating proteins can enable in-depth understanding of cellular physiology and may pave way towards developing novel methods of tackling disease progression. In this work, we exploit the existing wealth of data on phase separating proteins to develop sequence-based machine learning method for prediction of phase separating proteins. We use reduced alphabet schemes based on hydrophobicity and conformational similarity along with distributed representation of protein sequences and biochemical properties as input features to Support Vector Machine (SVM) and Random Forest (RF) machine learning algorithms. We used both curated and balanced dataset for building the models. RF trained on balanced dataset with hydropathy, conformational similarity embeddings and biochemical properties achieved accuracy of 97%. Our work highlights the use of conformational similarity, a feature that reflects amino acid flexibility, and hydrophobicity for predicting phase separating proteins. Use of such "interpretable" features obtained from the ever-growing knowledgebase of phase separation is likely to improve prediction performances further.

Download full-text PDF	Source
http://dx.doi.org/10.1109/TCBB.2022.3149310	DOI Listing

Publication Analysis

Top Keywords

phase separating

separating proteins

phase separation

conformational similarity

identification phase

reduced alphabet

phase

cellular physiology

machine learning

biochemical properties

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!