mold, a novel software to compile accurate and reliable DNA diagnoses for taxonomic descriptions.

Mol Ecol Resour

Institut Systématique Evolution Biodiversité (ISYEB), Muséum national d'Histoire naturelle, CNRS, Sorbonne Université, EPHE, Université des Antilles, Paris, France.

Published: July 2022

AI Article Synopsis

  • A new software program called "mold" has been developed to identify diagnostic nucleotide combinations (DNCs) that provide formal diagnoses for specific taxa.
  • The program also introduces a more reliable type of DNA diagnosis called "redundant DNC" (rDNC), which factors in unsampled genetic diversity, improving the reliability of taxonomic descriptions.

Article Abstract

DNA data are increasingly being used for phylogenetic inference, and taxon delimitation and identification, but scarcely for the formal description of taxa, despite their undisputable merits in taxonomy. The uncertainty regarding the robustness of DNA diagnoses, however, remains a major impediment to their use. We have developed a new program, mold, that identifies diagnostic nucleotide combinations (DNCs) in DNA sequence alignments for selected taxa, which can be used to provide formal diagnoses of these taxa. To test the robustness of DNA diagnoses, we carry out iterated haplotype subsampling for selected query species in published DNA data sets of varying complexity. We quantify the reliability of diagnosis by diagnosing each query subsample and then checking if this diagnosis remains valid against the entire data set. We demonstrate that widely used types of diagnostic DNA characters are often absent for a query taxon or are not sufficiently reliable. We thus propose a new type of DNA diagnosis, termed "redundant DNC" (or rDNC), which takes into account unsampled genetic diversity, and constitutes a much more reliable descriptor of a taxon. mold successfully retrieves rDNCs for all but two species in the analysed data sets, even in those comprising hundreds of species. mold shows unparalleled efficiency in large DNA data sets and is the only available software capable of compiling DNA diagnoses that suit predefined criteria of reliability.

Download full-text PDF

Source
http://dx.doi.org/10.1111/1755-0998.13590DOI Listing

Publication Analysis

Top Keywords

dna diagnoses
16
dna data
12
data sets
12
dna
10
robustness dna
8
diagnoses
5
data
5
mold
4
mold novel
4
novel software
4

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!