A Scan for Positively Selected Genes in the Genomes of Humans and Chimpanzees
If humans and chimpanzees share more than 98 percent of their DNA, and if most of that difference is silent — synonymous mutations that change a DNA letter without changing the protein — then the genes that actually made us human should be findable. You look for the places where protein-changing mutations piled up faster than silent ones. That excess is the molecular fingerprint of positive selection: evolution in a hurry. Nielsen and colleagues did exactly that scan across more than thirteen thousand human-chimp gene pairs. And the genes evolution was most in a hurry about? Not the ones building a bigger brain. Something else entirely. The method at the center of this work is built on a ratio called dN over dS. The numerator, dN, is the rate of nonsynonymous substitutions — mutations that change an amino acid. The denominator, dS, is the rate of synonymous substitutions — mutations that don't. Under neutral evolution, those two rates should be roughly equal, giving a ratio of about one. When dN over dS climbs above one, amino-acid-changing mutations are being fixed faster than silent ones, which means selection is actively pushing protein changes through. Nielsen and colleagues applied a one-sided likelihood ratio test of that neutral hypothesis across thirteen thousand seven hundred thirty-one orthologous gene pairs, after filtering out six thousand six hundred thirty that didn't meet quality thresholds.
For each pair, they used a codon-based model that accounted for transition-transversion rate bias and unequal codon frequencies. Because many gene pairs showed very little divergence, the standard chi-square approximation for the test statistic wasn't reliable, so they used simulations — conditioning on the observed number of nucleotide differences per gene — to determine the proper null distribution. For a gene of five hundred codons, their power simulations showed the test achieves more than eighty percent power when dN over dS reaches five. With that machinery in place, the question was what categories of genes would emerge. The field had reasonable expectations. Immune genes and sensory genes — those made sense. Pathogens evolve fast, so immune proteins need to keep up. Sensory receptors, especially olfactory ones, differ markedly between primates with different ecological niches. And indeed, immunity and defense topped the list, with a nominal p-value of zero under the PANTHER functional classification system. Sensory perception, including olfaction, also appeared prominently. These were the expected winners, and they delivered. Then things got strange.
Among the fifty genes showing the strongest evidence for positive selection, Nielsen and colleagues found a striking overrepresentation of tumor suppressors, apoptosis regulators, and cell-cycle control genes. Four putative tumor suppressors made the top fifty: the HYAL3 gene, the DFFA gene, the PEPP-2 gene, and the C16orf3 gene. Genes linked to tumor progression, like MMP26, appeared alongside apoptosis regulators including PPP1R15A, HSJ001348, TSARG1, and GZMH. The apoptosis inhibitor DFFA showed the strongest positive-selection signal among apoptosis genes and functions in Fas-mediated apoptosis — a pathway implicated in cancer control and in cell death during sperm production. The tissue-expression data made this stranger still. Genes with maximal expression in the testis were the only expression category to survive Bonferroni correction — the strict multiple-testing threshold — with two hundred forty-seven testis-expressed genes showing enrichment at a p-value of 0.0002. Many of the top candidates were testis- or sperm-specific: PRM1, USP26, C15orf2, PEPP-2, TCP11, HYAL3, and TSARG1. Genes involved in gametogenesis showed enrichment at a p-value of 0.0005, and spermatogenesis and motility at a p-value of 0.0037. And genes annotated as involved in inhibition of apoptosis showed excess evidence for positive selection at a p-value of 0.0047.
Meanwhile, brain genes showed almost nothing. Genes with maximal expression in the whole brain — eighty-three of them — had no excess of positively selected candidates, with a p-value of 0.965. Genes expressed in the brain at twice the level seen in blood actually trended toward avoidance of positive selection, at a p-value of 0.0002. Let that sit for a moment. The molecular scan that was supposed to tell us what made humans human found the weakest signals in the brain and the strongest signals in the testis and in genes that control cell death. The team's explanation for this pattern is a hypothesis about genomic conflict. Apoptosis — programmed cell death — is conspicuous in normal spermatogenesis, eliminating up to seventy-five percent of potential sperm cells before and after meiosis. That process exists for good reason at the organism level: it culls defective cells, controls output, and prevents genomic errors from propagating. But here's the conflict. Individual germ cells have their own evolutionary incentive. A mutation that lets a germ cell escape apoptosis, even slightly, will be favored at the cell level — that cell produces more descendants — even if it's harmful to the organism as a whole, because it increases cancer risk or disrupts somatic function.
Nielsen and colleagues argue that mutations acting after meiosis could produce segregation distortion: a variant that marginally increases a cell's odds of surviving postmeiotic culling can carry a large selective advantage at the gamete level. Compensatory changes in the apoptotic machinery would then be favored in response, setting up an arms race between selfish elements and the genes controlling cell death. That cycling pressure drives rapid evolution — which registers as positive selection — in the very genes that also happen to suppress tumors. The authors are careful to note that sperm competition and pathogen-driven selection can't be ruled out. But the overlap between Fas-mediated apoptosis in the germline and tumor-suppression pathways makes the conflict model a well-motivated account. The X chromosome adds another layer. Genes on the X show elevated positive selection at a p-value of 0.0049, and even after removing testis-expressed and spermatogenesis-annotated genes, the X remains enriched at a p-value of 0.0131. Males carry only one X, so X-linked mutations in males are immediately exposed to selection without a second copy masking their effects. That hemizygosity creates asymmetric selection pressure — a mechanism consistent with the genomic-conflict picture.
To check whether any of this was real rather than a statistical artifact of the comparison method, Nielsen and colleagues brought in a second, independent line of evidence. They sequenced the top fifty candidate genes in twenty Caucasian-American and nineteen African-American individuals and looked for the population-genetic fingerprint of a selective sweep: an excess of high-frequency derived nonsynonymous mutations. The logic is that when selection drives a beneficial variant toward fixation, it carries neighboring variants along for the ride, pushing them to high frequency in the population. Of the fifty genes, forty-six contained intraspecific polymorphism, yielding one hundred sixteen nonsynonymous polymorphisms and fifty-five synonymous ones. The nonsynonymous variants showed a clear excess at high frequencies. The synonymous ones didn't deviate in the same way. To make sure this wasn't an artifact of having selected genes based on their high dN over dS to begin with, the team ran one thousand simulated neutral datasets with the same distributions of divergence, mutation rates, and gene lengths, then re-selected the top fifty from each simulation. Those simulations showed the ascertainment procedure inflates the synonymous frequency spectrum — but essentially leaves the nonsynonymous spectrum untouched. So the excess of high-frequency derived nonsynonymous variants in the real data is not an artifact.
A comparison to the SeattleSNPs external database reinforced this: twenty-four of one hundred sixteen nonsynonymous polymorphisms in Nielsen and colleagues' data had derived-allele frequency above fifty percent, versus thirty-seven of three hundred sixty in the Seattle data — a significant difference at a p-value below 0.01. The team then fitted a Poisson Random Field model — a framework that treats the scaled selection coefficient S as two times population size times the per-mutation selection coefficient — to the nonsynonymous frequency spectrum. Their maximum-likelihood estimates for these fifty genes: approximately seventy-five percent of mutations are negatively selected, seventeen percent are effectively neutral, and eight percent are positively selected. Likelihood-ratio tests reject simpler models without positive selection at p-values of 0.0006 and 0.004 under different null specifications. Two lines of evidence — cross-species divergence and within-species polymorphism — pointing at the same fifty genes is a powerful convergence. That's the methodological contribution as much as any individual finding: combining comparative genomics with population genetics as a two-pronged screen, where each approach catches what the other might miss.
What the screen found is a picture of human-chimp molecular evolution that doesn't center on the brain. The coding-sequence signals are loudest in reproduction and cell death. That doesn't mean cognitive evolution wasn't real — it may simply mean that selection on cognition operated through regulatory changes, not protein-coding differences, which this kind of scan would miss entirely. But the genes that show the clearest molecular evidence of being driven by positive selection, the ones where evolution was in a hurry between us and our closest relatives, are the ones governing which sperm cells survive and which cells are allowed to keep dividing. Not the architecture of thought. The machinery of reproduction and death. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Diversity and scale: Genetic architecture of 2068 traits in the VA Million Veteran Program
- A Map of Recent Positive Selection in the Human Genome
- Positive feedback regulation between glycolysis and histone lactylation drives oncogenesis in pancreatic ductal adenocarcinoma
- Rapid phosphatidic acid accumulation in response to low temperature stress in Arabidopsis is generated through diacylglycerol kinase
- Organised Genome Dynamics in the Escherichia coli Species Results in Highly Diverse Adaptive Paths
- Mobile phone use and stress, sleep disturbances, and symptoms of depression among young adults - a prospective cohort study