More Recent Comments

Sunday, August 23, 2026

What is "alternate RNA decoding"?

I recently came across a paper in Nature on alternate RNA decoding. It looked interesting so I read the abstract. That didn't help. I was still confused about what the authors were assaying and whether it was anything other than translation or splicing errors. So I read the introduction to try and get a better understanding of alternate RNA decoding.

That didn't help. The rest of the paper wasn't any more helpful.

I'm going to reproduce the abstract and the relevant part of the introduction so you can judge for yourself whether I'm just being stupid or whether the authors should have been more clear about what they are doing. I think I'm reasonably well-informed about basic biochemistry and molecular biology principles and concepts but maybe I'm falling behind in my old age.

Nature is a prestigious journal and in the past (25 years ago) you could assume that any paper published in Nature was probably important. Today, I'm not so sure about that but maybe I'm being too critical.

There's another problem. I'm often finding that I can't understand the experiments and analyses that are described in current publications. This is especially true when there's a lot of bioinformatics. It all looks like black box science to me and I get worried when the authors announce a new revolutionary concept based on their results. Am I the only one who feels this way?

Tsour, S., Machne, R., Leduc, A., Widmer, S., Koo, E., Guez, J., Karczewski, K.J. and Slavov, N. (2026) Alternate RNA decoding results in stable and abundant proteins in mammals. Nature 656:506-515. [doi: 10.1038/s41586-026-10678-2]

Abstract

Amino acid substitutions may substantially alter protein stability and function. However, the contribution of substitutions that arise from alternate translation (deviations from the genetic code) is unknown. Here to address this issue, we analysed deep proteomic, transcriptomic and genomic data from more than 1,000 human samples, including 6 cancer types and 26 healthy human tissues. This global analysis identified 60,803 fragmentation spectra corresponding to 8,746 unique substitutions in proteins derived from 1,767 genes, including 1,955 confidently localized sites. Some substitutions were shared across samples, whereas others exhibited strong tissue-type and cancer specificity. Notably, products of alternate translation were more abundant than their canonical counterparts for hundreds of proteins, which suggests that there is sense-codon recoding. Recoded proteins included transcription factors, proteases, signalling proteins and proteins associated with neurodegeneration. Mechanisms that contribute to substitution abundance included protein stability, codon frequency, codon–anticodon mismatches and RNA modifications. We also characterized how alternatively translated proteoform ratios vary across protein domains, tissue types and cancers. These ratios were positively associated with intrinsically disordered regions and genetic polymorphisms in the gnomAD database, although the polymorphisms could not account for the substitutions. The sequence, relative abundance and the tissue specificity of alternatively translated proteins were conserved between humans and mice. These results demonstrate the contribution of alternate translation to the diversification of mammalian proteomes and its association with protein stability, tissue-specific proteomes and disease.

Introduction

Genetic mutations, RNA editing and ‘alternate translation’ (mRNA translation that deviates from the genetic code) may result in amino acid substitutions (AASs). Some substitutions that arise from mutations substantially change protein activity. For example, the V600E phosphomimetic substitution in BRAF destabilizes hydrophobic interactions and constitutively increases BRAF activity by up to 500-fold, which can induce tumorigenesis1. Similarly, the H1047R substitution in the catalytic subunit p110α of PI3K causes extensive cellular remodelling. Substitutions introduced by RNA editing also alter protein functions, such as the A-to-I edit in mRNA of the glutamate receptor.

In addition to arising from genetic mutations or editing of the RNA sequence, AASs may arise from alternate translation of the genetic code. Such alternate decoding can occur when all open reading frames (short or long, annotated or not annotated as protein coding) are translated, but it remains less characterized. AASs were historically detected on the basis of gel shifts of radioactively labelled proteins. Mass spectrometry (MS) substantially increased the power of detecting substitutions, and deep-learning predictions of peptide elution times and MS fragmentation spectra are facilitating the validation of noncanonical amino acid sequences. Translational deviations from the genetic code are considered rare and restricted to special cases.

If the rate of alternate translation is low, the abundance and significance of its protein products may also be low. However, the abundance of proteoforms with substitutions is determined not only by the rate of their synthesis but also by their degradation rate. Although most substitutions are likely to destabilize the substituted proteoform, some might stabilize them. By contrast, if a proteoform is stabilized by a substitution, it may accumulate to high abundance. Whether such protein stabilization contributes to an increase in levels of proteins synthesized by alternate RNA decoding remains unknown. To explore this question, we systematically identified and quantified AASs across human healthy and cancer tissues and mouse tissues. We identified, validated and characterized hundreds of abundant substitutions, including mechanisms that contribute to their origin and their impact on protein stability.

Here's the press release from Northeastern University [Expanding Proteome Diversity Through Alternate RNA Decoding.

BioE Professor Nikolai Slavov’s laboratory published research on “Alternate RNA decoding results in stable and abundant proteins in mammals” in Nature. Since the 1960s, the genetic code has been used to predict protein sequences from DNA and mRNA sequences. Slavov’s article demonstrates that these predictions miss thousands of protein sequences present in human tissues.

Across >1,000 human samples, we identified numerous abundant proteins whose amino acid sequences differ from those predicted by the genetic code.

These proteins are not rare translation byproducts. They accumulate to thousands of copies per cell. Some are more abundant than the proteins predicted by the genetic code from the same transcripts.

Their abundance reflects a combination of alternate RNA decoding mechanisms — including codon-anticodon mismatches, tRNA abundance, and RNA modifications — and selective stabilization of the resulting proteins. The last factor – protein stability – emerges as a major determinant of protein abundance across proteins, proteoforms, and cell types: https://slavovlab.net/research.htm#Proteostasis

Alternate RNA decoding is pervasive across functional groups of proteins, healthy and diseased tissues. It affects proteins playing key roles in neurodegeneration, and some alternately decoded proteins show strong enrichment in tumors compared to their surrounding tissues.

The findings reveal a layer of proteome diversity that is largely invisible to DNA and RNA sequences alone. Our knowledge of the proteome remains relatively limited: It is the next big Scientific Frontier.

This discovery has been a long and exhilarating journey with Shira Tsour and the Slavov Lab team. It started in 2019 and proceeded through many challenges and thrilling highs. A journey that has opened new perspectives that we long to explore!


15 comments :

Donald Forsdyke said...

They submitted in 2024 and it was not published until 2026, so the reviewing must have been problematic. Perhaps because many have been left behind in the key areas they mention - namely "intrinsically disordered proteins" and "RNA editing". My recent review of these topics in BioSystems (which likewise took a while to gain acceptance) might be helpful: doi: 10.1016/j.biosystems.2026.105900.

doi: 10.1016/j.biosystems.2026.105900

Anonymous said...

I found this article helpful: https://www.science.org/content/blog-post/when-variant-proteins-aren-t-actually-variant-ones

Anonymous said...

I think it simply means that a codon is misread and the wrong amino acid is incorportated into the protein. According to the paper it appears to happen quite frequently, to the degree that some codons in certain contexts seem to have a different amino acid associated with it than what the standard genetic code would imply (i.e. recoding).

SPARC said...

Being around doesn’t necessarily mean that any of these altered sequences has any function different from the respective wildtype peptide. They may even lack these functions or any other impact. Being enriched in a tumor doesn’t mean any such peptides has any function in tumor induction or progression necessarily. IMO being conserved in mice points to selection of the native sequence rather than conservation of co- or post-translational amino acids changes which may just represent unavoidable translational noise that may differ from tissue to tissue or change upon tumor induction.

Michael Tress said...

I have finally had chance to read the paper. It basically posits that changes at the nucleotide level during transcription are more frequent than expected and some of these end up being detectable as different proteins in proteomics experiments. Some of these changed proteins are even more abundant than the proteins predicted to be translated from the genomic sequence.

It is definitely an interesting addition since it adds yet another layer of complexity to the translation process. How much of it is under control and how much is due to transcriptional noise is left open. I do think that this sentence from the abstract "The sequence, relative abundance and the tissue specificity of alternatively translated proteins were conserved between humans and mice." is an exaggeration since only 5% of mouse cases coincided with human. They also missed a trick because the 55 cases that were found in both human and mouse experiments were not detailed and these would have been particularly interested to understand why they were conserved.

SPARC said...

Isn’t translation quite error prone? I can imagine that certain errors accumulate in cases where they don’t interfere with the cell’s quality control systems and don’t cause damage to the cells directly. I.e., such sequence changes could just be a consequence of the cellular systems and rather unavoidable noise than another level of regulation. In addition, wouldn’t one expect the same changes in mice and men if the underlying mechanisms leading to the accumulation of protein variants with altered peptide sequences are the same?

Michael Tress said...

In this paper they haven't tested that and they assume that the change in amino acid results from transcription. There is data that suggests that at least some of it likely to be due to transcription. I would think that translation could be error prone too, though. A lot more work would be needed to distinguish how much of this "decoding" is error and where it stems from.

SPARC said...

Maybe I am completely missreading the article but if you were right, I don’t see any reason why the paper had been published at all.
If it is about transcription, I wonder why the authors continuously refer to alternate translation, why they discuss the number of mismatches corresponding to alternate codon/anticodon binding, why they refer to translational fidelitys and why they begin the main part with “Genetic mutations, RNA editing and ‘alternate translation’ (mRNA translation that deviates from the genetic code) may result in amino acid substitutions (AASs).“ In addition why does the term „translation“ occur over 40 times and „transcription“ less than half as often?
IMO, they have some interesting results but struggle to show any biological relevance. Obviously, they are aware of this and try to sell it without causing an ENCODE 2.0 disaster by making any claims regarding function: “To avoid assuming functions or mechanisms, we chose the neutral phenomenological term alternate translation.” However, they sneak in the assumptions regarding functions in the next sentences which I addressed in my previous comment: “Although some SAAPs may merely reflect limits of translational fidelity and proteostasis, others may have evolved biological functions, as previously suggested, consistent with regulated sense-codon recoding discussed above. The high abundance, conservation across species and associations of some SAAPs with cancer and protein domains imply functional significance, and this possibility needs to be explored in future research.”

Larry Moran said...

@SPARC: I wasn't just published, it was published in Nature!

Brooklynette said...

Dr. John Timmer has commentary on a seemingly, perhaps related topic in an open access article, ""Automated prototyping of genetic codes" in Nature (8/26) at https://www.nature.com/articles/s41586-026-10949-y. Timmer writes that the authors, "designed a separate genetic code and used the alternative transfer RNAs to implement it. They then designed a messenger RNA that could be translated by both genetic codes, but would produce different proteins depending on which code was being used. They then put together a mixture of both populations of transfer RNAs, both populations of ribosomes, and all the chemicals needed to get translation to work.

Two different proteins were produced. So, both populations of ribosomes latched onto the messenger RNA but used different populations of transfer RNAs to make a protein using the messenger. And since the two populations implemented different genetic codes, the two populations of ribosomes made different proteins. This is really cool." See: https://arstechnica.com/science/2026/08/researchers-get-two-genetic-codes-to-work-at-the-same-time/

Of equal interest to readers of this blog, there is a "back and forth" in the comments section to Dr. Timmer's article, in which Larry's genome book and the concept of 90% junk DNA is defended by Dr Jay against some confusing and historically inaccurate arguments by a certain Xephrys

SPARC said...

To my best understanding Tsours and coworkers don’t say that alternate translation requires extra tRNAs. The closest they come to this is the discussion functionally altered amino-acyl-tRNAs which have been loaded with a wrong amino acid. However, IMO they prefer some translational error caused at the level of codon/anti-codon interaction as the likely explanation when they write “This analysis indicated a strong dependence between RAAS and the minimum number of mismatches required for the corresponding substitution. The lower probability of multiple base-pair mismatches was reflected in a much lower RAAS for an AAS requiring multiple mismatches.“

Michael Tress said...

@SPARC The authors probably call it alternative translation because they are detecting proteins. But it's not because they have evidence that the “recoding” happened during translation. Despite all the work they did, all we know is that these changes happen at some point between the DNA sequence and the detection of the protein sequence in the cell. So they could happen during either transcription or translation, but also, for example, during RNA processing, via post-translational modifications (not impossible even though they have taken them into account) or in the detection process (mis-identification). And possibly all of them. All of them would be my guess.

I suppose Nature has published this study because it is large-scale and a deviation from the expected towards further apparent complexity. These sort of papers are popular at the moment.

Anonymous said...

If the altered protein sequences have been caused by transcription errors or incorrect RNA processing it doesn’t make sense to talk about alternate RNA decoding because this would be a change of the coding sequence not of the code itself. The encoded amino acid would fit to its codon present in the mRNA. In addition, post-transcriptional modifications of RNAs, incorporation of amino acids like selenocysteine and pyrrolysine which use stop codons, and amino acids created by post-translational modifications are well established concepts with known underlying processes which actually don’t change the genetic code. My impression was that the authors tried to rule out such causes. If one thinks they did not I don’t see any justification to publish the paper at all.
I wonder if it makes biological sense to add yet another level of regulation/complexity. IMO it may make sense in certain cellular pathways if the underlying mechanisms (i.e., not just the presence of a fraction of an altered protein) can be established. However, I am not an expert in evolution theory but I wonder if and how selection would work on alternately translated proteins.
One last thought: Would a similar paper claiming that mis- or alteranatively folded cellular proteins form another level of regulation just because they can be identified have been published?

SPARC said...

The last comment was mine. Unfortunately, I couldn't log in when I posted it.

Anonymous said...

Yes, you are too old.