Interpreting the CTCF-mediated sequence grammar of genome folding with AkitaV2.

PLoS Comput Biol

Department of Quantitative and Computational Biology, University of Southern California, Los Angeles, California, United States of America.

Published: February 2025

Interphase mammalian genomes are folded in 3D with complex locus-specific patterns that impact gene regulation. CTCF (CCCTC-binding factor) is a key architectural protein that binds specific DNA sites, halts cohesin-mediated loop extrusion, and enables long-range chromatin interactions. There are hundreds of thousands of annotated CTCF-binding sites in mammalian genomes; disruptions of some result in distinct phenotypes, while others have no visible effect. Despite their importance, the determinants of which CTCF sites are necessary for genome folding and gene regulation remain unclear. Here, we update and utilize Akita, a convolutional neural network model, to extract the sequence preferences and grammar of CTCF contributing to genome folding. Our analyses of individual CTCF sites reveal four predictions: (i) only a small fraction of genomic sites are impactful; (ii) impact is highly dependent on sequences flanking the core CTCF binding motif; (iii) core and flanking nucleotides contribute largely additively to the overall impact of a site; (iv) sites created as combinations of different core and flanking sequences have impacts proportional to the product of their average impacts, i.e. they are broadly compatible. Our analysis of collections of CTCF sites make two predictions for multi-motif grammar: (i) insulation strength depends on the number of CTCF sites within a cluster, and (ii) pattern formation is governed by the orientation and spacing of these sites, rather than any inherent specialization of the CTCF motifs themselves. In sum, we present a framework for using neural network models to probe the sequences instructing genome folding and provide a number of predictions to guide future experimental inquiries.

Download full-text PDF

Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC11828424PMC
http://dx.doi.org/10.1371/journal.pcbi.1012824DOI Listing

Publication Analysis

Top Keywords

genome folding
16
ctcf sites
16
sites
9
mammalian genomes
8
gene regulation
8
ctcf
8
neural network
8
core flanking
8
interpreting ctcf-mediated
4
ctcf-mediated sequence
4

Similar Publications

Induction of MASH-like pathogenesis in the Nwd1 mouse liver.

Commun Biol

March 2025

Laboratory for Molecular Neurobiology, Faculty of Human Sciences, Waseda University, Tokorozawa, Saitama, Japan.

Endoplasmic reticulum (ER) stores Ca and plays crucial roles in protein folding, lipid transfer, and it's perturbations trigger an ER stress. In the liver, chronic ER stress is involved in the pathogenesis of metabolic dysfunction-associated steatotic liver disease (MASLD) and metabolic dysfunction-associated steatohepatitis (MASH). Dysfunction of sarco/endoplasmic reticulum calcium ATPase (SERCA2), a key regulator of Ca transport from the cytosol to ER, is associated with the induction of ER stress and lipid droplet formation.

View Article and Find Full Text PDF

All the members of the phylum Cnidaria are characterized by the production of venom in specialized structures, the nematocysts. Venom of jellyfish (Medusozoa) and sea anemones (Anthozoa) has been investigated since the 1970s, revealing a remarkable molecular diversity. Specifically, sea anemones harbour a rich repertoire of neurotoxic peptides, some of which have been developed in drug leads.

View Article and Find Full Text PDF

Protein synthesis by ribosomes produces functional proteins but also serves diverse regulatory functions, which depend on the coding amino acid sequences. Certain nascent peptides interact with the ribosome exit tunnel to arrest translation and modulate themselves or the expression of downstream genes. However, a comprehensive understanding of the mechanisms of such ribosome stalling and its regulation remains elusive.

View Article and Find Full Text PDF

Herein, we reported mutations in five DNA Damage Repair (DDR) i.e., TP53, ATR, ATM, CHEK1 and CHEK2 involved in OSCC using NG-WES and their analysis using bioinformatics tools.

View Article and Find Full Text PDF

DNA G-quadruplexes (G4s) are non-B-form DNA secondary structures that threaten genome stability by impeding DNA replication. To elucidate how G4s induce replication fork arrest, we characterized fork collisions with preformed G4s in the parental DNA using reconstituted yeast and human replisomes. We demonstrate that a single G4 in the leading strand template is sufficient to stall replisomes by arresting the CMG helicase.

View Article and Find Full Text PDF

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!