The number of species with high-quality genome sequences continues to increase, in part due to the scaling up of multiple large-scale biodiversity sequencing projects. While the need to annotate genic sequences in these genomes is widely acknowledged, the parallel need to annotate transposable element (TE) sequences that have been shown to alter genome architecture, rewire gene regulatory networks, and contribute to the evolution of host traits is becoming ever more evident. However, accurate genome-wide annotation of TE sequences is still technically challenging. Several de novo TE identification tools are now available, but manual curation of the libraries produced by these tools is needed to generate high-quality genome annotations. Manual curation is time-consuming, and thus impractical for large-scale genomic studies, and lacks reproducibility. In this work, we present the Manual Curator Helper tool MCHelper, which automates the TE library curation process. By leveraging MCHelper's fully automated mode with the outputs from three de novo TE identification tools, RepeatModeler2, EDTA, and REPET, in the fruit fly, rice, hooded crow, zebrafish, maize, and human, we show a substantial improvement in the quality of the TE libraries and genome annotations. MCHelper libraries are less redundant, with up to 65% reduction in the number of consensus sequences, have up to 11.4% fewer false positive sequences, and up to ∼48% fewer "unclassified/unknown" TE consensus sequences. Genome-wide TE annotations are also improved, including larger unfragmented insertions. Moreover, MCHelper is an easy-to-install and easy-to-use tool.

Download full-text PDF

Source
http://dx.doi.org/10.1101/gr.278821.123DOI Listing

Publication Analysis

Top Keywords

transposable element
8
high-quality genome
8
novo identification
8
identification tools
8
manual curation
8
genome annotations
8
consensus sequences
8
sequences
7
mchelper
4
mchelper automatically
4

Similar Publications

Background: Anorexia nervosa (AN) is a polygenic, severe metabopsychiatric disorder with poorly understood aetiology. Eight significant loci have been identified by genome-wide association studies (GWAS) and single nucleotide polymorphism (SNP)-based heritability was estimated to be ~ 11-17, yet causal variants remain elusive. It is therefore important to define the full spectrum of genetic variants in the wider regions surrounding these significantly associated loci.

View Article and Find Full Text PDF

The chromatin remodeling factor OsINO80 promotes H3K27me3 and H3K9me2 deposition and maintains TE silencing in rice.

Nat Commun

December 2024

State Key Laboratory of Genetic Engineering, Collaborative Innovation Center of Genetics and Development, Department of Biochemistry, Institute of Plant Biology, School of Life Sciences, Fudan University, Shanghai, PR China.

The INO80 chromatin remodeling complex plays a critical role in shaping the dynamic chromatin environment. The diverse functions of the evolutionarily conserved INO80 complex have been widely reported. However, the role of INO80 in modulating the histone variant H2A.

View Article and Find Full Text PDF

Background/aim: Lung cancer, a predominant contributor to cancer mortality, is characterized by diverse etiological factors, including tobacco smoking and genetic susceptibilities. Despite advancements, particularly in nonsmall-cell lung cancer (NSCLC), therapeutic options for lung squamous cell carcinoma (LUSC) are limited. Transposable elements (TEs) and their regulatory proteins, such as tigger transposable element derived (TIGD) family proteins, have been implicated in cancer development.

View Article and Find Full Text PDF

DNA methylation is an essential epigenetic mechanism for regulation of gene expression, through which many physiological (X-chromosome inactivation, genetic imprinting, chromatin structure and miRNA regulation, genome defense, silencing of transposable elements) and pathological processes (cancer and repetitive sequences-associated diseases) are regulated. Nanopore sequencing has emerged as a novel technique that can analyze long strands of DNA (long-read sequencing) without chemically treating the DNA. Interestingly, nanopore sequencing can also extract epigenetic status of the nucleotides (including both 5-Methylcytosine and 5-hydroxyMethylcytosine), and a large variety of bioinformatic tools have been developed for improving its detection properties.

View Article and Find Full Text PDF

Characterization and comparative analysis of antimicrobial resistance in from hospital and municipal wastewater treatment plants.

J Water Health

December 2024

Department of Microbiology, Kasturba Medical College, Manipal, Manipal Academy of Higher Education, Manipal, 576104, Karnataka, India; Center for Antimicrobial Resistance and Education (CARE), Manipal Academy of Higher Education, Manipal, 576104, Karnataka, India E-mail:

The spread of antimicrobial resistance (AMR) poses global health threats, with wastewater treatment plants (WWTPs) as hotspots for its development. Horizontal gene transfer facilitates acquisition of resistance genes, particularly through integrons in . Our study investigates isolates from hospital and municipal WWTPs, focusing on integrons, their temporal correlation and phenotypic and molecular characterization of AMR.

View Article and Find Full Text PDF

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!