The genome sequence of the Mamavirus, a new Acanthamoeba polyphaga mimivirus strain, is reported. With 1,191,693 nt in length and 1,023 predicted protein-coding genes, the Mamavirus has the largest genome among the known viruses. The genomes of the Mamavirus and the previously described Mimivirus are highly similar in both the protein-coding genes and the intergenic regions.
View Article and Find Full Text PDFAccurate inference of orthologous genes is a pre-requisite for most comparative genomics studies, and is also important for functional annotation of new genomes. Identification of orthologous gene sets typically involves phylogenetic tree analysis, heuristic algorithms based on sequence conservation, synteny analysis, or some combination of these approaches. The most direct tree-based methods typically rely on the comparison of an individual gene tree with a species tree.
View Article and Find Full Text PDFBackground: Accurate estimation of the divergence time of the extant eukaryotes is a fundamentally important but extremely difficult problem owing primarily to gross violations of the molecular clock at long evolutionary distances and the lack of appropriate calibration points close to the date of interest. These difficulties are intrinsic to the dating of ancient divergence events and are reflected in the large discrepancies between estimates obtained with different approaches. Estimates of the age of Last Eukaryotic Common Ancestor (LECA) vary approximately twofold, from ~1,100 million years ago (Mya) to ~2,300 Mya.
View Article and Find Full Text PDFMTH1203, a β-CASP metallo-β-lactamase family nuclease from the archaeon Methanothermobacter thermautotrophicus, was identified as a putative nuclease that might contribute to RNA processing. The crystal structure of MTH1203 reveals that, in addition to the metallo-β-lactamase nuclease and the β-CASP domains, it contains two contiguous KH domains that are unique to MTH1203 and its orthologs. RNA-binding experiments indicate that MTH1203 preferentially binds U-rich sequences with a dissociation constant in the micromolar range.
View Article and Find Full Text PDFThe CRISPR-Cas (clustered regularly interspaced short palindromic repeats-CRISPR-associated proteins) modules are adaptive immunity systems that are present in many archaea and bacteria. These defence systems are encoded by operons that have an extraordinarily diverse architecture and a high rate of evolution for both the cas genes and the unique spacer content. Here, we provide an updated analysis of the evolutionary relationships between CRISPR-Cas systems and Cas proteins.
View Article and Find Full Text PDFThe widespread exchange of genes among prokaryotes, known as horizontal gene transfer (HGT), is often considered to "uproot" the Tree of Life (TOL). Indeed, it is by now fully clear that genes in general possess different evolutionary histories. However, the possibility remains that the TOL concept can be reformulated and remain valid as a statistical central trend in the phylogenetic "Forest of Life" (FOL).
View Article and Find Full Text PDFThe division of labor between template and catalyst is a fundamental property of all living systems: DNA stores genetic information whereas proteins function as catalysts. The RNA world hypothesis, however, posits that, at the earlier stages of evolution, RNA acted as both template and catalyst. Why would such division of labor evolve in the RNA world? We investigated the evolution of DNA-like molecules, i.
View Article and Find Full Text PDFWe describe the draft genome of the microcrustacean Daphnia pulex, which is only 200 megabases and contains at least 30,907 genes. The high gene count is a consequence of an elevated rate of gene duplication resulting in tandem gene clusters. More than a third of Daphnia's genes have no detectable homologs in any other available proteome, and the most amplified gene families are specific to the Daphnia lineage.
View Article and Find Full Text PDFPlants possess two myosin classes, VIII and XI. The myosins XI are implicated in organelle transport, filamentous actin organization, and cell and plant growth. Due to the large size of myosin gene families, knowledge of these molecular motors remains patchy.
View Article and Find Full Text PDFClustered Regularly Interspaced Short Palindromic Repeats (CRISPRs) and the associated proteins (Cas) comprise a system of adaptive immunity against viruses and plasmids in prokaryotes. Cas1 is a CRISPR-associated protein that is common to all CRISPR-containing prokaryotes but its function remains obscure. Here we show that the purified Cas1 protein of Escherichia coli (YgbT) exhibits nuclease activity against single-stranded and branched DNAs including Holliday junctions, replication forks and 5'-flaps.
View Article and Find Full Text PDFThe highly conserved Kinase, Endopeptidase and Other Proteins of small Size (KEOPS)/Endopeptidase-like and Kinase associated to transcribed Chromatin (EKC) protein complex has been implicated in transcription, telomere maintenance and chromosome segregation, but its exact function remains unknown. The complex consists of five proteins, Kinase-Associated Endopeptidase (Kae1), a highly conserved protein present in bacteria, archaea and eukaryotes, a kinase (Bud32) and three additional small polypeptides. We showed that the complex is required for a universal tRNA modification, threonyl carbamoyl adenosine (t6A), found in all tRNAs that pair with ANN codons in mRNA.
View Article and Find Full Text PDFBackground: It is common belief that all cellular life forms on earth have a common origin. This view is supported by the universality of the genetic code and the universal conservation of multiple genes, particularly those that encode key components of the translation system. A remarkable recent study claims to provide a formal, homology independent test of the Universal Common Ancestry hypothesis by comparing the ability of a common-ancestry model and a multiple-ancestry model to predict sequences of universally conserved proteins.
View Article and Find Full Text PDFRegulation of gene expression during infection of the thermophilic bacterium Thermus thermophilus HB8 with the bacteriophage P23-45 was investigated. Macroarray analysis revealed host transcription shut-off and identified three temporal classes of phage genes; early, middle and late. Primer extension experiments revealed that the 5' ends of P23-45 early transcripts are preceded by a common sequence motif that likely defines early viral promoters.
View Article and Find Full Text PDFThe first congress on Viruses of Microbes took place at the Institut Pasteur in Paris, France, on 21-25 June 2010. The advances in genomics and metagenomics reported at this meeting reveal striking and unexpected complexity of the virus world. Viruses, in particular viruses that infect prokaryotes and unicellular eukaryotes, are emerging as the most abundant class of biological entities on earth and a major evolutionary and geochemical force.
View Article and Find Full Text PDFActa Crystallogr Sect F Struct Biol Cryst Commun
October 2010
New distinct versions of known protein folds provide a powerful means of protein-function prediction that complements sequence and genomic context analysis. These structures do not supplant direct biochemical experiments, but are indispensable for the complete characterization of proteins.
View Article and Find Full Text PDFThe recent discovery of protein modification by SAMPs, ubiquitin-like (Ubl) proteins from the archaeon Haloferax volcanii, prompted a comprehensive comparative-genomic analysis of archaeal Ubl protein genes and the genes for enzymes thought to be functionally associated with Ubl proteins. This analysis showed that most archaea encode members of two major groups of Ubl proteins with the β-grasp fold, the ThiS and MoaD families, and indicated that the ThiS family genes are rarely linked to genes for thiamine or Mo/W cofactor metabolism enzymes but instead are most often associated with genes for enzymes of tRNA modification. Therefore it is hypothesized that the ancestral function of the archaeal Ubl proteins is sulfur insertion into modified nucleotides in tRNAs, an activity analogous to that of the URM1 protein in eukaryotes.
View Article and Find Full Text PDFPhylogenetic trees of individual genes of prokaryotes (archaea and bacteria) generally have different topologies, largely owing to extensive horizontal gene transfer (HGT), suggesting that the Tree of Life (TOL) should be replaced by a "net of life" as the paradigm of prokaryote evolution. However, trees remain the natural representation of the histories of individual genes given the fundamentally bifurcating process of gene replication. Therefore, although no single tree can fully represent the evolution of prokaryote genomes, the complete picture of evolution will necessarily combine trees and nets.
View Article and Find Full Text PDFThe majority of mammalian genes produce multiple transcripts resulting from alternative splicing (AS) and/or alternative transcription initiation (ATI) and alternative transcription termination (ATT). Comparative analysis of the number of alternative nucleotides, isoforms, and introns per locus in genes with different types of alternative events suggests that ATI and ATT contribute to the diversity of human and mouse transcriptome even more than AS. There is a strong negative correlation between AS and ATI in 5' untranslated regions (UTRs) and AS in coding sequences (CDSs) but an even stronger positive correlation between AS in CDSs and ATT in 3' UTRs.
View Article and Find Full Text PDFNat Rev Microbiol
October 2010
Recently a novel cell division system comprised of homologues of eukaryotic ESCRT-III (endosomal sorting complex required for transport III) proteins was discovered in the hyperthermophilic crenarchaeote Sulfolobus acidocaldarius. On the basis of this discovery, we undertook a comparative genomic analysis of the machineries for cell division and vesicle formation in Archaea. Archaea possess at least three distinct membrane remodelling systems: the FtsZ-based bacterial-type system, the ESCRT-III-based eukaryote-like system and a putative novel system that uses an archaeal actin-related protein.
View Article and Find Full Text PDFThe rapidly accumulating genome sequence data allow researchers to address fundamental biological questions that were not even asked just a few years ago. A major problem in genomics is the widening gap between the rapid progress in genome sequencing and the comparatively slow progress in the functional characterization of sequenced genomes. Here we discuss two key questions of genome biology: whether we need more genomes, and how deep is our understanding of biology based on genomic analysis.
View Article and Find Full Text PDFA long-standing assumption in evolutionary biology is that the evolution rate of protein-coding genes depends, largely, on specific constraints that affect the function of the given protein. However, recent research in evolutionary systems biology revealed unexpected, significant correlations between evolution rate and characteristics of genes or proteins that are not directly related to specific protein functions, such as expression level and protein-protein interactions. The strongest connections were consistently detected between protein sequence evolution rate and the expression level of the respective gene.
View Article and Find Full Text PDFMost of the archaea and numerous bacteria possess an elaborate system of adaptive immunity to mobile genetic elements known as the CRISPR (clustered regularly interspaced short palindromic repeats)-associated system (CRISPR-Cas), which consists of arrays of short repeats interspersed with unique DNA spacers and adjacent operons encompassing CRISPR-associated (cas) genes with predicted and, in some cases, experimentally validated nuclease, helicase, and polymerase activities. The system functions by integrating fragments of alien DNA between the repeats and employing their transcripts to degrade the DNA of the respective invading elements via an RNA interference-like mechanism. The CRISPR-Cas system is a case of apparent Lamarckian inheritance.
View Article and Find Full Text PDFIntervirology
September 2010
Background/aims: The nucleo-cytoplasmic large DNA viruses (NCLDV) constitute an apparently monophyletic group that consists of 6 families of viruses infecting a broad variety of eukaryotes. A comprehensive genome comparison and maximum-likelihood reconstruction of NCLDV evolution reveal a set of approximately 50 conserved genes that can be tentatively mapped to the genome of the common ancestor of this class of eukaryotic viruses. We address the origins and evolution of NCLDV.
View Article and Find Full Text PDFMultiple constraints variously affect different parts of the genomes of diverse life forms. The selective pressures that shape the evolution of viral, archaeal, bacterial and eukaryotic genomes differ markedly, even among relatively closely related animal and bacterial lineages; by contrast, constraints affecting protein evolution seem to be more universal. The constraints that shape the evolution of genomes and phenomes are complemented by the plasticity and robustness of genome architecture, expression and regulation.
View Article and Find Full Text PDFBackground: The arylalkylamine N-acetyltransferase (AANAT) family is divided into structurally distinct vertebrate and non-vertebrate groups. Expression of vertebrate AANATs is limited primarily to the pineal gland and retina, where it plays a role in controlling the circadian rhythm in melatonin synthesis. Based on the role melatonin plays in biological timing, AANAT has been given the moniker "the Timezyme".
View Article and Find Full Text PDF