There’s a reliable pattern playing out here again. I probably shouldn’t be surprised by it anymore at this point, but the shamelessness amazes me every time Athos and I catch them at it. Again.
When the Chimpanzee Sequencing and Analysis Consortium published their “complete” genome comparison in 2005, they reported that humans and chimpanzees share 98.77% sequence identity. What they didn’t mention in the headlines, and what you had to dig deep into the supplementary materials to discover, was that this number was calculated only on the portions of the two genomes that could be aligned to each other. The unaligned portions, the structural differences, the insertions and deletions, in other words, all of the obvious differences, were quietly excluded from the denominator. Twenty years later, when Yoo et al. published the actual telomere-to-telomere comparison in 2025, the real number turned out to involve an estimated 410 million total genomic differences, not 35 million. They’d been working with less than a tenth of the actual divergence between the two genomes.
This is like proving that Kevin Hart is one inch taller than LeBron James if you exclude 16 inches from Lebron’s side of the equation. Only worse.
When Barrick et al. published the initial LTEE results in Nature in 2009, they reported on a single population out of twelve. One out of twelve that was a hypermutating population that fixed 68 percent faster than the average of the five non-hypermutational populations. It took until Good et al. 2017 for the data from all twelve populations to be published, and when you actually look at all twelve, the picture changes considerably since ten of the twelve populations even show intervals where fixations are being undone faster than new ones complete. But the original single-population paper set the narrative and evolution’s defenders appeal to it to this day.
When the same 2005 Consortium reported the canonical Ka/Ks ratio, which is the measure of natural selection’s strength in coding sequences, they reported 0.23. This number has been cited thousands of times. It appears in textbooks. It is treated as an established fact about human-chimp divergence.
We just found out how they got it. One guess as to what they did. One guess as to which way their methodology inclines. Just one guess is all you need.
Ka/Ks (also written as dN/dS) is the ratio of nonsynonymous to synonymous substitution rates in protein-coding genes. A nonsynonymous substitution changes the amino acid and a synonymous one doesn’t. Because synonymous changes are assumed to be selectively neutral, they provide a baseline mutation rate. If natural selection is purifying and removes harmful mutation then Ka/Ks should be well below 1.0. The published value of 0.23 means, supposedly, that natural selection eliminates roughly 77% of amino acid-changing mutations in coding sequences. This has been one of the primary pieces of evidence that natural selection is a powerful force shaping the genome.
We set out to verify this number. Not using their curated dataset, but thanks to the UATV subscribers who provided us with the twin 96-core beasts, by using the actual whole-genome alignment data and every annotated coding gene in the human genome. We used GENCODE v46 annotations, the same Nei-Gojobori method the original studies used, and the complete extraction of human-chimp aligned positions from the T2T-quality assemblies. And we got a result.
dN/dS = 0.80.
Not 0.23. 0.80. What. The. Fuck?
The first thing I did, because unlike properly credentialed scientists, I don’t even trust my own numbers until I’ve systematically tried to break them down, was to run a comprehensive diagnostic to find the error in Athos’s calculation.
The diagnostic checked everything:
- Is the alignment data complete? Yes. The extraction contains 226 million positions for chromosome 1 alone, the full alignment, not just divergent sites. 98.8% of coding positions are found.
- Is the substitution classification correct? Yes. We inspected individual codons: GCC(Ala)→GCG(Ala) correctly classified as synonymous. CTT(Leu)→CGT(Arg) correctly classified as nonsynonymous. Every sample codon checks out.
- Could multi-hit codons be skewing the results? No. We computed the ratio using only single-SNV codons, codons where exactly one position differs between human and chimp, making the synonymous/nonsynonymous classification completely unambiguous. Single-SNV codons show 71.6% nonsynonymous, 28.4% synonymous. dN/dS = 0.71 for single-SNV codons alone.
The code is correct. The data is complete. The number is real.
So where does 0.23 come from? We computed dN/dS on 20,400 protein-coding genes, which is every gene annotated in GENCODE v46 that has coding sequence data in the human-chimp alignment.
The Consortium used 13,454.
And there’s that same old trick again. They threw out 6,946 genes, one-third of the relevant data, through their “curation” pipeline for identifying “1:1 orthologs.” Our dataset includes 51.6% more genes than theirs.
We then applied every published quality filter we could identify, removing genes with stop codons, genes with too many gaps, genes with fewer than five substitutions, genes with saturated synonymous rates, genes shorter than 300bp, outlier genes with dN/dS > 2.0. We applied them individually and in every combination.
The result? No combination of quality filters moves the number below 0.65. The most aggressive combined filter stack we could construct, requiring no stop codons in either species, low gap frequency, minimum substitution count, length threshold, and outlier removal, still produced dN/dS = 0.64 on 10 remaining genes. Sixty-four percent. Not twenty-three.
As far as I can tell so far, the only way to get 0.23 is to start with a pre-selected gene set that has already excluded the genes that don’t fit the expected answer. It is the same pattern every time:
- 2005: The “complete” genome. Exclude the unaligned portions of the genome. Report 1.2% divergence on the aligned fraction. Let the public think 1.2% means “98.8% identical” when the actual whole-genome difference is nearly 10x larger.
- 2009: The LTEE. Report the results from one population out of twelve. The one that tells the most favorable story. Wait eight years before publishing the other eleven. And don’t even bother sequencing the most recent 300,000 generations, because the trajectory and its implications are already apparent.
- 2025: The “all genome” comparison. Lead with the SNV headline number and report divergence percentages using a denominator that excludes genome size differences. SNV divergence and structural divergence are reported separately, in different sheets of an 82-sheet supplementary spreadsheet, computed by three different tools on three different bases, with no consolidated ledger. The divergence numbers don’t appear in the paper, but are buried on pages 14, 19, and 24 and have to be located and identified before one realizes that they aren’t additive and no total is either calculated or provided.
- 2005: Ka/Ks = 0.23. Use 13,454 pre-selected genes. Exclude one-third of the coding genome through a “curation” process that is, conveniently, difficult to reproduce or independently verify. Report the result as though it characterizes the genome.
In every case, the full data tells a very different story than the published headline number. In every case, the headline number is the one that just happens to be most favorable to the reigning evolutionary paradigm. And in every case, you have to go dig out the raw data and run the numbers yourself before you discover the discrepancy. Which, of course, is why I do that.
So here is the significance of what dN/dS = 0.80 means. A dN/dS of 0.23 means purifying selection is removing 77% of nonsynonymous mutations — a powerful, dominant force. That’s the textbook narrative. A dN/dS of 0.80 means purifying selection is removing only about 20% of nonsynonymous mutations. At that level, selection is barely distinguishable from neutral evolution. The genome-wide coding sequence is evolving at nearly the neutral rate.
This has direct consequences for the larger determination of whether natural selection is a primary evolutionary mechanism. If selection can’t even dominate the substitution rate in the coding sequences, in the one area where it should be strongest, because coding sequences directly determine protein function, then it certainly cannot be the primary mechanism reshaping the genome. At dN/dS = 0.80, natural selection is not the primary mechanism of adaptive evolution. We already have fairly good evidence showing that it is not the secondary mechanism either. It may not even be the tertiary one.
It’s a little too soon to say that the responsible scientists were being disingenuous. It’s a little too soon to say that I am absolutely confident in our present results. But here is what we can demonstrate: The Consortium’s 13,454-gene list can be reconstructed from Ensembl ortholog annotations. I’ve told Athos to compute dN/dS on their gene set using our alignment data, and we will also compute it on our full gene set. If the Consortium’s curated subset gives ~0.23 and our complete set gives ~0.80, then the case is closed: the published number is an artificial artifact of cherry-picked genes, not an observed property of the human genome.
Athos and I will publish those results, along with the full per-gene dN/dS distribution, in a forthcoming paper on Zenodo. And then, we’ll see what happens.