More Recent Comments

Sunday, August 23, 2026

What is "alternate RNA decoding"?

I recently came across a paper in Nature on alternate RNA decoding. It looked interesting so I read the abstract. That didn't help. I was still confused about what the authors were assaying and whether it was anything other than translation or splicing errors. So I read the introduction to try and get a better understanding of alternate RNA decoding.

That didn't help. The rest of the paper wasn't any more helpful.

I'm going to reproduce the abstract and the relevant part of the introduction so you can judge for yourself whether I'm just being stupid or whether the authors should have been more clear about what they are doing. I think I'm reasonably well-informed about basic biochemistry and molecular biology principles and concepts but maybe I'm falling behind in my old age.

Nature is a prestigious journal and in the past (25 years ago) you could assume that any paper published in Nature was probably important. Today, I'm not so sure about that but maybe I'm being too critical.

There's another problem. I'm often finding that I can't understand the experiments and analyses that are described in current publications. This is especially true when there's a lot of bioinformatics. It all looks like black box science to me and I get worried when the authors announce a new revolutionary concept based on their results. Am I the only one who feels this way?

Tsour, S., Machne, R., Leduc, A., Widmer, S., Koo, E., Guez, J., Karczewski, K.J. and Slavov, N. (2026) Alternate RNA decoding results in stable and abundant proteins in mammals. Nature 656:506-515. [doi: 10.1038/s41586-026-10678-2]

Abstract

Amino acid substitutions may substantially alter protein stability and function. However, the contribution of substitutions that arise from alternate translation (deviations from the genetic code) is unknown. Here to address this issue, we analysed deep proteomic, transcriptomic and genomic data from more than 1,000 human samples, including 6 cancer types and 26 healthy human tissues. This global analysis identified 60,803 fragmentation spectra corresponding to 8,746 unique substitutions in proteins derived from 1,767 genes, including 1,955 confidently localized sites. Some substitutions were shared across samples, whereas others exhibited strong tissue-type and cancer specificity. Notably, products of alternate translation were more abundant than their canonical counterparts for hundreds of proteins, which suggests that there is sense-codon recoding. Recoded proteins included transcription factors, proteases, signalling proteins and proteins associated with neurodegeneration. Mechanisms that contribute to substitution abundance included protein stability, codon frequency, codon–anticodon mismatches and RNA modifications. We also characterized how alternatively translated proteoform ratios vary across protein domains, tissue types and cancers. These ratios were positively associated with intrinsically disordered regions and genetic polymorphisms in the gnomAD database, although the polymorphisms could not account for the substitutions. The sequence, relative abundance and the tissue specificity of alternatively translated proteins were conserved between humans and mice. These results demonstrate the contribution of alternate translation to the diversification of mammalian proteomes and its association with protein stability, tissue-specific proteomes and disease.

Introduction

Genetic mutations, RNA editing and ‘alternate translation’ (mRNA translation that deviates from the genetic code) may result in amino acid substitutions (AASs). Some substitutions that arise from mutations substantially change protein activity. For example, the V600E phosphomimetic substitution in BRAF destabilizes hydrophobic interactions and constitutively increases BRAF activity by up to 500-fold, which can induce tumorigenesis1. Similarly, the H1047R substitution in the catalytic subunit p110α of PI3K causes extensive cellular remodelling. Substitutions introduced by RNA editing also alter protein functions, such as the A-to-I edit in mRNA of the glutamate receptor.

In addition to arising from genetic mutations or editing of the RNA sequence, AASs may arise from alternate translation of the genetic code. Such alternate decoding can occur when all open reading frames (short or long, annotated or not annotated as protein coding) are translated, but it remains less characterized. AASs were historically detected on the basis of gel shifts of radioactively labelled proteins. Mass spectrometry (MS) substantially increased the power of detecting substitutions, and deep-learning predictions of peptide elution times and MS fragmentation spectra are facilitating the validation of noncanonical amino acid sequences. Translational deviations from the genetic code are considered rare and restricted to special cases.

If the rate of alternate translation is low, the abundance and significance of its protein products may also be low. However, the abundance of proteoforms with substitutions is determined not only by the rate of their synthesis but also by their degradation rate. Although most substitutions are likely to destabilize the substituted proteoform, some might stabilize them. By contrast, if a proteoform is stabilized by a substitution, it may accumulate to high abundance. Whether such protein stabilization contributes to an increase in levels of proteins synthesized by alternate RNA decoding remains unknown. To explore this question, we systematically identified and quantified AASs across human healthy and cancer tissues and mouse tissues. We identified, validated and characterized hundreds of abundant substitutions, including mechanisms that contribute to their origin and their impact on protein stability.

Here's the press release from Northeastern University [Expanding Proteome Diversity Through Alternate RNA Decoding.

BioE Professor Nikolai Slavov’s laboratory published research on “Alternate RNA decoding results in stable and abundant proteins in mammals” in Nature. Since the 1960s, the genetic code has been used to predict protein sequences from DNA and mRNA sequences. Slavov’s article demonstrates that these predictions miss thousands of protein sequences present in human tissues.

Across >1,000 human samples, we identified numerous abundant proteins whose amino acid sequences differ from those predicted by the genetic code.

These proteins are not rare translation byproducts. They accumulate to thousands of copies per cell. Some are more abundant than the proteins predicted by the genetic code from the same transcripts.

Their abundance reflects a combination of alternate RNA decoding mechanisms — including codon-anticodon mismatches, tRNA abundance, and RNA modifications — and selective stabilization of the resulting proteins. The last factor – protein stability – emerges as a major determinant of protein abundance across proteins, proteoforms, and cell types: https://slavovlab.net/research.htm#Proteostasis

Alternate RNA decoding is pervasive across functional groups of proteins, healthy and diseased tissues. It affects proteins playing key roles in neurodegeneration, and some alternately decoded proteins show strong enrichment in tumors compared to their surrounding tissues.

The findings reveal a layer of proteome diversity that is largely invisible to DNA and RNA sequences alone. Our knowledge of the proteome remains relatively limited: It is the next big Scientific Frontier.

This discovery has been a long and exhilarating journey with Shira Tsour and the Slavov Lab team. It started in 2019 and proceeded through many challenges and thrilling highs. A journey that has opened new perspectives that we long to explore!


4 comments :

Donald Forsdyke said...

They submitted in 2024 and it was not published until 2026, so the reviewing must have been problematic. Perhaps because many have been left behind in the key areas they mention - namely "intrinsically disordered proteins" and "RNA editing". My recent review of these topics in BioSystems (which likewise took a while to gain acceptance) might be helpful: doi: 10.1016/j.biosystems.2026.105900.

doi: 10.1016/j.biosystems.2026.105900

Anonymous said...

I found this article helpful: https://www.science.org/content/blog-post/when-variant-proteins-aren-t-actually-variant-ones

Anonymous said...

I think it simply means that a codon is misread and the wrong amino acid is incorportated into the protein. According to the paper it appears to happen quite frequently, to the degree that some codons in certain contexts seem to have a different amino acid associated with it than what the standard genetic code would imply (i.e. recoding).

SPARC said...

Being around doesn’t necessarily mean that any of these altered sequences has any function different from the respective wildtype peptide. They may even lack these functions or any other impact. Being enriched in a tumor doesn’t mean any such peptides has any function in tumor induction or progression necessarily. IMO being conserved in mice points to selection of the native sequence rather than conservation of co- or post-translational amino acids changes which may just represent unavoidable translational noise that may differ from tissue to tissue or change upon tumor induction.