Large scale sequence alignment via efficient inference in generative models.

Mihir Mongia Chengze Shen Arash Gholami Davoodi Guillaume Marçais Hosein Mohimani

Sci Rep

School Computer Science, Carnegie Mellon University, Pittsburgh, USA.

Published: May 2023

Finding alignments between millions of reads and genome sequences is crucial in computational biology. Since the standard alignment algorithm has a large computational cost, heuristics have been developed to speed up this task. Though orders of magnitude faster, these methods lack theoretical guarantees and often have low sensitivity especially when reads have many insertions, deletions, and mismatches relative to the genome. Here we develop a theoretically principled and efficient algorithm that has high sensitivity across a wide range of insertion, deletion, and mutation rates. We frame sequence alignment as an inference problem in a probabilistic model. Given a reference database of reads and a query read, we find the match that maximizes a log-likelihood ratio of a reference read and query read being generated jointly from a probabilistic model versus independent models. The brute force solution to this problem computes joint and independent probabilities between each query and reference pair, and its complexity grows linearly with database size. We introduce a bucketing strategy where reads with higher log-likelihood ratio are mapped to the same bucket with high probability. Experimental results show that our method is more accurate than the state-of-the-art approaches in aligning long-reads from Pacific Bioscience sequencers to genome sequences.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC10160065	PMC
http://dx.doi.org/10.1038/s41598-023-34257-x	DOI Listing

Publication Analysis

Top Keywords

sequence alignment

genome sequences

probabilistic model

query read

log-likelihood ratio

large scale

scale sequence

alignment efficient

efficient inference

inference generative

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!