More Recent Comments

Wednesday, September 30, 2026

Dark matter and junk DNA

It's been almost three years since I created a blog post on What is the "dark matter of the genome"?. I noted that the term was popularized by the philosopher Evelyn Fox Keller who was opposed to the concept of junk DNA. She preferred to think that most of the genome was full of functional sequences that hadn't yet been discovered. The term "dark matter" was her way of emphasizing that most of the genome was a mystery but it wasn't junk.

Since then, the term "dark matter" has been increasingly used by scientists, especially those in the field of genomics. I think it's for the same reasons that Evelyn Fox Keller described. These scientists spend a lot of time and effort analyzing entire genomes for biochemical activity of various sorts and they don't want to admit that they are mostly looking at junk DNA. That's why I summarized my position like this ...

The phrase "dark matter of the genome" is used by scientists who are skeptical of junk DNA so they want to convey the impression that most of the genome consists of important DNA whose function is just waiting to be discovered. Not surprisingly, the term is often used by researchers who are looking for funding and investors to support their efforts to use the latest technology to discover this mysterious function that has eluded other scientists for over 50 years.

I was reminded of this controversy recently when I came across a number of papers published by The Codebook Consortium. This is a large international collaboration of scientists who are trying to figure out the possible functions of a set of mysterious transcription factors. Many of the authors are based at the University of Toronto where I spent most of my scientific career.

Here's one of the papers that illustrates the length scientists will go to in order to avoid saying the words "junk DNA."

Razavi, R. et al. (2026) Extensive binding of uncharacterized human transcription factors to genomic dark matter. Nat. Commun. 17:7770. [doi: 10.1038/s41467-026-75376-z]

The functional impact of a large portion of the human genome known as “dark matter DNA”, which is composed mainly of repeat sequences, remains unknown. The genome also encodes many putative and poorly characterized transcription factors. Here, we determine genomic binding locations of 166 poorly characterized human transcription factors in living cells. Nearly half of them associate strongly with known regulatory regions such as promoters and enhancers, frequently co-localizing with each other at conserved motif matches. The other half often associate with genomic dark matter, however, at largely non-overlapping (i.e., unique) sites, via intrinsic sequence recognition. Fifty-four of the latter half, which we term dark transcription factors, mainly bind within regions of closed chromatin, with each recognizing a unique set of repeat sequences. The dark transcription factors include many KZNFs, which are known to bind and silence transposable elements, and other transcription factors with apparent repressive functions. Others may be pioneer transcription factors. For example, we find that induction of TPRX1, a known regulator of zygotic preimplantation, leads to chromatin opening at many of its binding sites in the dark matter genome.

I think the best way to illustrate the difference between my perspective and that of the authors is to quote the first paragraph of their paper and then re-write it from my point of view. Keep in mind that my view is that 10% of the human genome is functional and 90% is junk. The functional elements include protein-coding DNA, regulatory sequences, centromeres, telomeres, origins of replication, and chromatin organizing sites (I call them SARs). So far, we know the function of about 8% of the genome but there's still an additional 2% that's under purifying selection but has no known function. This 2% is the real "dark matter" but I see no reason to use that term.

Here's the first paragraph from the paper.

For decades, the term dark matter, borrowed from astronomy and cosmology, has been applied to describe the excess of seemingly non-functional regions of many genomes. These regions are dominated by repetitive sequences, such as transposable elements (TEs) and endogenous retroelements (EREs), which largely account for the extreme differences in genome size across eukaryotes. There are many potential roles for genomic dark matter, but it is also conceivable that it is mainly inert or simply provides spacing. An initial proposal for the function of the dark matter was gene regulation, and indeed, while some gene deserts appear to be dispensable, others contain transcriptional enhancers. The function of genomic dark matter in gene regulation is also suggested by the existence of transcription factors (TFs) that are capable of binding TE and ERE sequences.

And here's how I would have written it.

For decades, we've known that only 10% of the human genome is functional and the rest is junk DNA. The junk DNA regions include introns and they consist largely of pseudogenes, fossil transposons, and fossil virus sequences that are evolving at the neutral rate. The functional sequences are subject to purifying selection and they contain regulatory elements that control gene expression. These regulatory elements, comprising a small fraction of the functional genome, consist largely of sites that bind transcription factors. Most of these binding sites consist of short DNA sequences, typically about eight base pairs in length, and these occur frequently, by chance, in the 90% of the genome that is junk DNA. For example, a specific eight base pair sequence will occur about 188,000 times in the human genome. It is often difficult to distinguish between these fortuitous, non-functional, binding sites and genuine regulatory sites that lie within the few percent of the genome that has yet to be characterized.

I think my version is more scientifically accurate and lays out the problem in a much better manner than the authors' version. What do you think?

There are many ways of detecting transcription factor binding sites in cells and it turns out that only a fraction of the potential binding sites are detected. This is probably because most of the potential binding sites are in closed chromatin domains that are not accessible to the transcription factors. [Open and closed chromatin domains (and epigenetics)].

The authors show that about 80% of the transcription factor binding sites are not located in known enhancer or promoter sequences and that's consistent with the idea that they are bound to junk DNA. If you read the discussion carefully you will discover that they authors recognize this possibility even though they never use the words "junk DNA."

Establishing physiological functions, if any, for individual TF binding sites is a long-standing and difficult problem in regulatory genomics.... Biochemical functions of TFs may be more straightforward to identify: ... Biochemical function does not equate to physiological purpose (and resultant selective pressure), however. A particular challenge with TFs is that the level of binding site turnover observed on evolutionary timescales requires that binding sites arise at random, many of which are likely irrelevant for gene regulation or reproductive fitness (at least initially), despite being biochemically functional. By this reasoning, we expect that many biochemically verified, direct TF binding sites should be physiologically non-functional, and indeed, we find that, overall, most TOP sites are not conserved, even for promoter TFs. Lack of conservation cannot be taken as a lack of biological purpose; nonetheless, conserved TOP sites would seem most likely to yield interpretable results in targeted laboratory studies. More generally, the Codebook TOP and CTOP catalogue will provide a rich resource for future efforts in examining genome function, e.g., employing techniques similar to GTEx71 to relate variants at the TOP sites to chromatin and gene expression.

Note the use of "biochemical function" instead of "biochemical activity"! That seems a bit disingenuous to me given all the controversy surrounding the ENCODE debacle in 2012.

Note also the use of two typical ENCODE excuses. The first is a reminder that the lack of conservation doesn't mean lack of function. (This is expressed more clearly in other parts of the paper.) The idea here is that most transcription factor binding sites are not conserved and that's because regulatory sites turn over rapidly. Thus, there may be hundreds of regulatory sites that have recently evolved in the human lineage and are not present in other apes or primates. Not only is that highly unlikely, the excuse also ignores the data on lack of purifying selection.

The second excuse is more subtle and it's related to the defense offered by ENCODE researchers back in 2012-2014. It's true, they say, that much of their data consists of mapping irrelevant, non-functional, sites in junk DNA but this database may be useful ("rich resource") for "future efforts in examining genome function." That hasn't worked out in the past ten years.

We still haven't heard any of these genomics labs come up with a solid number for the amount of DNA devoted to regulating gene expression. (I think it's about 0.2%.)


The top figure is an AI-generated image created by Gemini. The bottom one is from the paper.

1 comment :

Anonymous said...

Larry, how can you ignore the strong evidence for the role of the dark genome in snipes and woozles?