Summary: Bioinformatics applications increasingly rely on ad hoc disk storage of k-mer sets, e.g. for de Bruijn graphs or alignment indexes. Here, we introduce the K-mer File Format as a general lossless framework for storing and manipulating k-mer sets, realizing space savings of 3-5× compared to other formats, and bringing interoperability across tools.

Availability And Implementation: Format specification, C++/Rust API, tools: https://github.com/Kmer-File-Format/.

Supplementary Information: Supplementary data are available at Bioinformatics online.

Download full-text PDF

Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC9477520PMC
http://dx.doi.org/10.1093/bioinformatics/btac528DOI Listing

Publication Analysis

Top Keywords

k-mer file
8
file format
8
k-mer sets
8
k-mer
4
format standardized
4
standardized compact
4
compact disk
4
disk representation
4
representation sets
4
sets k-mers
4

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!