AI Article Synopsis

  • - The text discusses the challenges of inferring past demographic histories in population genetics due to limitations in current methods, despite the availability of many complete genomes.
  • - It introduces a new framework called approximate Bayesian computation based on the random forest algorithm (ABC-RF), which uses a statistical approach that efficiently summarizes genomic data through the full distribution of segregating sites.
  • - The authors tested the accuracy of this new method using simulated and real data, focusing on models related to the migration of modern humans and the evolutionary relationships of orangutans.

Article Abstract

Inferring past demographic histories is crucial in population genetics, and the amount of complete genomes now available should in principle facilitate this inference. In practice, however, the available inferential methods suffer from severe limitations. Although hundreds complete genomes can be simultaneously analysed, complex demographic processes can easily exceed computational constraints, and the procedures to evaluate the reliability of the estimates contribute to increase the computational effort. Here we present an approximate Bayesian computation framework based on the random forest algorithm (ABC-RF), to infer complex past population processes using complete genomes. To this aim, we propose to summarize the data by the full genomic distribution of the four mutually exclusive categories of segregating sites (FDSS), a statistic fast to compute from unphased genome data and that does not require the ancestral state of alleles to be known. We constructed an efficient ABC pipeline and tested how accurately it allows one to recognize the true model among models of increasing complexity, using simulated data and taking into account different sampling strategies in terms of number of individuals analysed, number and size of the genetic loci considered. We also compared the FDSS with the unfolded and folded site frequency spectrum (SFS), and for these statistics we highlighted the experimental conditions maximizing the inferential power of the ABC-RF procedure. We finally analysed real data sets, testing models on the dispersal of anatomically modern humans out of Africa and exploring the evolutionary relationships of the three species of Orangutan inhabiting Borneo and Sumatra.

Download full-text PDF

Source
http://dx.doi.org/10.1111/1755-0998.13263DOI Listing

Publication Analysis

Top Keywords

complete genomes
12
random forest
8
approximate bayesian
8
bayesian computation
8
data
5
distinguishing complex
4
complex evolutionary
4
evolutionary models
4
models unphased
4
unphased whole-genome
4

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!