Whole genome sequencing of bacteria is important to enable strain classification. Using entire genomes as an input to machine learning (ML) models would allow rapid classification of strains while using information from multiple genetic elements. We developed a "bag-of-words" approach to encode, using SentencePiece or k-mer tokenization, entire bacterial genomes and analyze these with ML. Initial model selection identified SentencePiece with 8,000 and 32,000 words as the best approach for genome tokenization. We then classified in genomes the capsule B group genotype with 99.6% accuracy and the multifactor invasive phenotype with 90.2% accuracy, in an independent test set. Subsequently, in silico knockouts of 2,808 genes confirmed that the ML model predictions aligned with our current understanding of the underlying biology. To our knowledge, this is the first ML method using entire bacterial genomes to classify strains and identify genes considered relevant by the classifier.
Download full-text PDF |
Source |
---|---|
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC10910294 | PMC |
http://dx.doi.org/10.1016/j.isci.2024.109257 | DOI Listing |
Viruses
November 2024
Department of Infectious Diseases, Molecular Virology, Section Virus-Host Interactions, Heidelberg University, 69120 Heidelberg, Germany.
The study of hepatitis C virus (HCV) replication in cell culture is mainly based on cloned viral isolates requiring adaptation for efficient replication in Huh7 hepatoma cells. The analysis of wild-type (WT) isolates was enabled by the expression of SEC14L2 and by inhibitors targeting deleterious host factors. Here, we aimed to optimize cell culture models to allow infection with HCV from patient sera.
View Article and Find Full Text PDFViruses
November 2024
Faculty of Medical and Health Sciences, Tel Aviv University, Tel Aviv 6997801, Israel.
In this study, we introduce a novel approach that integrates interpretability techniques from both traditional machine learning (ML) and deep neural networks (DNN) to quantify feature importance using global and local interpretation methods. Our method bridges the gap between interpretable ML models and powerful deep learning (DL) architectures, providing comprehensive insights into the key drivers behind model predictions, especially in detecting outliers within medical data. We applied this method to analyze COVID-19 pandemic data from 2020, yielding intriguing insights.
View Article and Find Full Text PDFSensors (Basel)
December 2024
Department of Information Technology, College of Computers and Information Technology, Taif University, P.O. Box 11099, Taif 21944, Saudi Arabia.
[...
View Article and Find Full Text PDFSensors (Basel)
December 2024
Department of Computer Science, School of Computing and Engineering, University of Huddersfield, Queensgate, Huddersfield HD1 3DH, UK.
Climate change caused by greenhouse gas (GHG) emissions is an escalating global issue, with the transportation sector being a significant contributor, accounting for approximately a quarter of all energy-related GHG emissions. In the transportation sector, vehicle emissions testing is a key part of ensuring compliance with environmental regulations. The Vehicle Certification Agency (VCA) of the UK plays a pivotal role in certifying vehicles for compliance with emissions and safety standards.
View Article and Find Full Text PDFEnter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!