Circular RNAs Are the Predominant Transcript Isoform from Hundreds of Human Genes in Diverse Cell Types
Why would a cell bother to make an RNA where the exons appear out of order? That question has lingered on the edges of RNA biology for decades. Hints had accumulated.
Electron microscopy suggested circular RNA existed in eukaryotes more than thirty years before this work, and a handful of mammalian examples had been documented, including the mouse SRY gene. Its RNA in adult testes exists primarily as a circular product that appears not to be translated. But these were just curiosities, with low abundance that required sensitive nested polymerase chain reaction to detect, often dismissed as transcriptional noise.
The consensus was that canonical splicing preserved genomic exon order and that anything else was rare and probably inconsequential.
Salzman and colleagues decided to ask that question at scale. Using deep paired-end RNA sequencing of ribosomal-RNA depleted total RNA from human cells, they looked for transcripts that could not be explained by conventional splicing. Because the libraries used roughly three hundred to five hundred base pair inserts and eighty base pair reads per end, they could detect reads spanning exon-exon junctions and ask a simple but powerful question: are these exons joined in the right order or the wrong one?
The answer was striking. In diagnostic bone marrow from five children with hyperdiploid B-precursor acute lymphoblastic leukemia, they found more than one thousand two hundred thirty-two genes with at least one junctional read consistent with exon scrambling — a downstream exon joined to an upstream one, the opposite of what genomic order would predict. But leukemia cells can have all kinds of genomic rearrangements.
So Salzman and colleagues sequenced three normal leukocyte populations from a single individual: CD19-positive B cells, CD34-positive hematopoietic stem cells, and neutrophils. Across those three normal cell types, they found sequence evidence for scrambled exons comprising at least ten percent of transcripts from each of more than eight hundred genes. This was not a cancer artifact. It was everywhere.
The scale deepened. Across all samples analyzed, the team identified two thousand seven hundred forty-eight distinct transcript isoforms with scrambled exon order. By a conservative counting method, two hundred twenty-nine, two hundred seven, and one hundred twenty-two scrambled transcripts — representing four hundred eighty-one distinct genes — were expressed at levels comparable to canonical linear isoforms in the B cell, stem cell, and neutrophil samples, respectively.
By a slightly looser criterion, eight hundred eighty distinct genes had scrambled isoforms at least ten percent as abundant as the canonical form. For some genes, the numbers were even more dramatic: KIAA0182, MAN1A2, and CCDC126 had scrambled isoforms that accounted for more than half of all transcripts. The predominant form of the gene's output was the non-canonical one.
Now, that finding demands scrutiny. Scrambled exon reads could arise from genomic tandem duplications, from rare trans-splicing between separate molecules, or from sequencing artifacts. Salzman and colleagues confronted each of these alternatives directly, and their experimental design reads like a systematic attempt to disprove their own result.
On the mapping side, they used paired-end read statistics to ask whether the scrambled junctions could be explained by linear transcripts. For each candidate gene, they computed the probability of seeing zero reads mapping outside the inferred exon interval — essentially, the probability that all paired reads landed inside a region consistent with a circular molecule rather than a longer linear one. This probability is the empirical insert-length distribution evaluated at the inferred circle length, raised to the power of the number of supporting junctional reads.
Across genes, the data showed a systematic underrepresentation of reads that would support a linear explanation. Only twenty-three genes — fewer than two percent of scrambled cases — had direct read evidence consistent with a linear scrambled transcript.
The biochemical test was RNase R treatment. RNase R is a three-prime to five-prime exonuclease that degrades linear RNA from its ends but cannot touch a circular molecule because circles have no ends. Salzman and colleagues treated HeLa total RNA with zero, three, ten, or one hundred units of RNase R per microgram of RNA, then ran validation polymerase chain reactions.
Canonical linear transcripts were sensitive — they disappeared with increasing enzyme. Transcripts with scrambled exons were resistant. That result is inconsistent with a reverse-transcription template-switching artifact, which would not survive RNase R treatment.
Northern blotting for MAN1A2 using a four hundred eighty-one base probe hybridized to a band matching the predicted circular isoform. Reverse-transcription polymerase chain reaction with no-template and no-reverse-transcription controls produced no products, ruling out contamination. Sanger sequencing confirmed the junction sequences directly.
Multiple orthogonal lines of evidence, each capable of disproving the circular hypothesis — and none of them did.
Salzman and colleagues also considered whether these circles arise via lariat intermediates — a known byproduct of canonical splicing where the intron loops into a lasso shape before being released. That mechanism predicts specific complementary linear transcripts. They found evidence for those predicted transcripts in only thirteen of five hundred seventy-six distinct scrambled isoforms, just two point two percent.
The lariat pathway wasn't the explanation. The most consistent interpretation was back-splicing: the spliceosome joining a downstream splice donor to an upstream splice acceptor, looping the RNA back on itself to form a covalently closed circle.
With the finding established in leukemia and leukocytes, the team extended the analysis outward. They surveyed diverse normal and malignant human cell types and found the same pattern. Then they turned to a publicly available mouse RNA sequencing dataset and applied the same analytical pipeline.
More than one thousand mouse genes showed evidence of exon scrambling. The phenomenon crosses species, which strongly suggests it is a conserved aspect of gene expression rather than a human peculiarity.
Subcellular localization added another dimension. Salzman and colleagues fractionated HeLa cells into nuclear and cytoplasmic compartments, verified fractionation quality using XIST as a nuclear control and by checking ribosomal RNA patterns on denaturing gels, then assayed circular isoforms by quantitative polymerase chain reaction. Most circular isoforms were more enriched in the cytoplasmic fraction than their canonical linear counterparts.
That matters because the cytoplasm is where translation happens, where RNA-binding proteins operate, and where regulatory interactions play out. Circular RNAs are not being sequestered in the nucleus and degraded. They are exported. They are present where cytoplasmic functions would occur.
Scrambled junction reads were also roughly ten times more frequent in poly-A depleted RNA than in poly-A selected RNA from HeLa and H9 cells. Most messenger RNAs carry a poly-A tail — circular RNAs, lacking free ends, cannot. That enrichment pattern is exactly what you would predict if these molecules are genuine circles.
On function, Salzman and colleagues are careful. They document abundance. They document cytoplasmic localization.
They note that circular ANRIL had previously been associated with INK4 and ARF expression as well as atherosclerosis risk, and that the SRY circular RNA appears not to be translated. But they do not claim specific molecular roles for the hundreds of new circles they found. The honest position is that the function of most of these molecules is unknown.
They could be regulatory. They could influence the abundance of linear isoforms by competing for splicing factors. They could have independent roles in the cytoplasm. The paper opens those questions; it does not close them.
What it does close is a simpler question: are circular RNAs common in human cells? The answer is yes — for hundreds of genes, emphatically yes. And that has immediate practical consequences.
Most RNA sequencing analysis pipelines assume that transcripts are linear. They are built around the expectation of canonical exon order. If a gene's dominant isoform is circular, a standard pipeline will either miscount it, misassign it, or miss it entirely.
Gene expression atlases, differential expression studies, transcript quantification — all of it was built on an assumption that Salzman and colleagues showed is wrong for a substantial fraction of the human genome.
The transcriptome has a layer that was invisible because the tools weren't looking for it. Now we know it's there.
This lecture was created by ennepō.
Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field.
Read when you can. Listen when you want to.
Related lectures
- Revised Estimates for the Number of Human and Bacteria Cells in the Body
- Dynamic regulation of genome-wide pre-mRNA splicing and stress tolerance by the Sm-like protein LSm5 in Arabidopsis
- The Pervasive Effects of an Antibiotic on the Human Gut Microbiota, as Revealed by Deep 16S rRNA Sequencing
- Population Structure and Eigenanalysis
- Accurate prediction of protein structures and interactions using a three-track neural network
- DNA methylation age of human tissues and cell types