A Map of Recent Positive Selection in the Human Genome

Benjamin F. Voight, Sridhar Kudaravalli, Xiaoquan Wen, Jonathan K. PritchardView original
OverviewBalancedriya_rao voice
A geneticist stares at a stretch of DNA that looks almost identical across hundreds of people — too identical, suspiciously so. That uniformity isn't noise. It's a signal. It means something swept through that population recently, carrying one version of that stretch of chromosome along for the ride, so fast that recombination never had time to break it apart. That's what positive selection looks like from inside the genome. Reading that signal systematically across the entire human genome at once is what Voight, Kudaravalli, Wen, and Pritchard did. The core problem they faced is a timing one. When a beneficial allele is genuinely new — still rising and not yet fixed across the population — many of the classic methods for detecting selection go quiet. Those methods look for signatures left behind after selection is finished, such as long-term divergence between species or alleles that have already swept to near-universal frequency. However, the variants under the most recent selection are still segregating. They haven't won yet. Candidate-gene studies, where researchers already suspected something interesting, had picked up a handful of examples: lactase persistence, a salt-sensitivity variant in the CYP3A5 gene, and signals near genes involved in malaria resistance. But those were hypothesis-driven, one locus at a time. What Voight and colleagues wanted was a map — a genome-wide, first-generation map of evolution still in progress. The resource that made it possible was the International HapMap Project, which provided roughly eight hundred thousand single nucleotide polymorphisms — single-letter differences in the genetic code — assayed in two hundred nine unrelated individuals drawn from three continental populations: Yoruba from Nigeria, northern and western Europeans, and East Asians. That density of markers, spread across the whole genome, finally made a systematic search feasible. The method they developed to do this is called iHS — the integrated haplotype score — and it works by reading haplotype length as a proxy for time. A haplotype is a chunk of chromosome inherited together. The longer and more uniform it is, the less time recombination has had to chop it up. When a beneficial allele rises quickly, the haplotype it rides on stays long. The key observable is something called EHH, or extended haplotype homozygosity, which measures the probability that two chromosomes carrying the same allele are genetically identical across increasingly distant sites. EHH starts at one right at the focal variant and decays toward zero as you move outward. Voight and colleagues integrated the area under that decay curve separately for the ancestral version of each variant and the derived version — the newer mutation. Then, they took the log ratio of those two integrated areas and standardized it by allele frequency. In plain terms, iHS is asking whether the derived allele sits on an unusually long haplotype compared to what you would expect for a variant at that frequency. Large negative values flag derived alleles with suspiciously long haplotypes — the signature of a sweep in progress. They computed iHS for over six hundred seventy thousand single nucleotide polymorphisms in Europeans, over seven hundred thousand in Yoruba, and over six hundred twenty thousand in East Asians, using a high-resolution recombination map built from linkage disequilibrium data. The sanity check is reassuring: the lactase region in Europeans, LCT, lights up brilliantly. That stretch spans about two thousand eight hundred kilobases, contains three hundred fifty-one out of five hundred ninety-four single nucleotide polymorphisms with an absolute iHS above two, and reaches a maximum iHS of four point nine — one of the most extreme values in the entire European dataset. That's a variant the field already knew was under strong recent selection, associated with the ability to digest dairy into adulthood, almost certainly driven by the spread of pastoralism. The method finds it cleanly. Now scale that up to the whole genome. What Voight and colleagues found is that signals of recent positive selection are widespread across all three populations — but are mostly local. They scanned in one hundred kilobase windows, flagging windows in the top one percent of their empirical distribution as candidate sweep regions. Most of those windows were population-specific. That's the headline: adaptation over this timescale has been regional, shaped by the particular environments and pressures each population encountered. However, there was a statistically clear excess of windows shared between two or even all three populations, more than chance would predict. This suggests some sweeps are responses to challenges that humans faced broadly. Then comes the result that flips an earlier assumption. Low-resolution genome scans had sometimes suggested that recent selection was relatively sparse in sub-Saharan Africans. Voight and colleagues find the opposite: by several measures, some of their strongest signals come from the Yoruba sample. The Yoruba sweep regions are narrower — average genetic spans of zero point thirty-two centimorgans, versus zero point fifty-two centimorgans in East Asians and Europeans — which likely explains why earlier coarser analyses missed them. Narrower regions mean older or stronger sweeps, not weaker ones. The age estimates, derived from haplotype spans assuming a twenty-five-year generation time, put average Yoruba sweeps at roughly ten thousand eight hundred years old versus about six thousand six hundred years in non-Africans. These events are recent by evolutionary standards, falling mainly within the Holocene — the last ten thousand years, the era of agriculture, changing diets, new pathogens, and dramatic shifts in population density and geography. The longest individual haplotypes point to the most extreme events. Near the GBA gene in East Asians, a selected haplotype extends one point thirty-nine centimorgans. Near NKX2-2 in Europeans, it extends one point twenty-five centimorgans. On chromosome five p fifteen in Yoruba, it extends zero point ninety-seven centimorgans in what the authors describe as a gene desert — the genetic target there remains unknown. This brings up the functional biology. Voight and colleagues cataloged what kinds of genes sit near the strongest selection signals, and several themes emerge clearly. Metabolism is prominent: LCT for lactase in Europeans, the SI gene for sucrose processing in East Asians, MAN2A1 in Yoruba and East Asians, and multiple genes tied to fat metabolism including LEPR — the leptin receptor, involved in regulating body fat mass — in East Asians, and NCOA1 in Yoruba. Detoxification shows up strongly too, particularly in Europeans, where cytochrome P450 genes — enzymes that break down a wide range of compounds — are significantly enriched. This includes CYP3A5, CYP2E1, and CYP1A2, as well as four genes in a CYP cluster on chromosome one. Pigmentation genes appear in Europeans: OCA2, MYO5A, DTNBP1, and TYRP1. Neurological candidates include two microcephaly-associated genes, CDK5RAP2 in Yoruba and CENPJ in Europeans and East Asians, plus the serotonin transporter SLC6A4 in Europeans and East Asians, and SNTG1 across all three populations. The genome-wide Gene Ontology analysis turned up enrichment in chemosensory perception, with p-values of zero point zero zero zero six in Europeans and zero point zero zero zero four in Yoruba. There is olfaction at similar significance, carbohydrate metabolism in East Asians at a p-value of zero point zero zero two, electron transport in Europeans at a p-value of zero point zero zero two, and reproduction-related categories with p-values around zero point zero zero three to zero point zero zero four. An unexpected finding in Yoruba is that the X chromosome is strikingly enriched, with fifteen genes mapped to top one percent windows in the Yoruba sample versus only six in Europeans and two in East Asians. Voight and colleagues are careful to state the central caveat plainly: for most of these signals, the actual phenotype under selection is unknown. Even when a candidate gene is identifiable, what it was selected for often remains unclear. These are starting points, not finished explanations. That caveat coexists with a forward-looking argument. If a variant has been rising under strong selection — simulations put the selection coefficients in the range of zero point zero one to zero point zero four — it must be producing real differences in fitness, which almost certainly means real differences in phenotype. And real phenotypic differences are exactly what genetic mapping studies need. A few of the identified regions already have established medical relevance: CYP3A5 and salt-sensitive hypertension, the alcohol dehydrogenase cluster ADH and alcoholism susceptibility, and the chromosomal inversion on seventeen q twenty-one and fertility. On that basis, Voight and colleagues built a practical tool: genome-wide iHS scores for HapMap single nucleotide polymorphisms and a set of tag SNPs designed to capture the top approximately two hundred fifty selection signals in each population. The logic is that wherever selection has recently pushed an allele toward high frequency, that region is worth examining for associations with complex traits. The map that Voight, Kudaravalli, Wen, and Pritchard produced is a first-generation chart of where the human genome has been actively shaped in the last several thousand years. Most signals are local. Some are shared. The strongest, by some measures, come from Africa. The genes involved touch metabolism, detoxification, sensory processing, reproduction, and brain development. The timing puts most events squarely in the agricultural era. What drove each sweep, in most cases, remains to be determined. But the map exists now — and it shows, unmistakably, that human evolution did not stop. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

A geneticist stares at a stretch of DNA that looks almost identical across hundreds of people — too identical, suspiciously so. That uniformity isn't noise. It's a signal. It means something swept through that population recently, carrying one version of that stretch of chromosome along for the ride, so fast that recombination never had time to break it apart. That's what positive selection looks like from inside the genome. Reading that signal systematically across the entire human genome at once is what Voight, Kudaravalli, Wen, and Pritchard did. The core problem they faced is a timing one. When a beneficial allele is genuinely new — still rising and not yet fixed across the population — many of the classic methods for detecting selection go quiet. Those methods look for signatures left behind after selection is finished, such as long-term divergence between species or alleles that have already swept to near-universal frequency. However, the variants under the most recent selection are still segregating. They haven't won yet. Candidate-gene studies, where researchers already suspected something interesting, had picked up a handful of examples: lactase persistence, a salt-sensitivity variant in the CYP3A5 gene, and signals near genes involved in malaria resistance. But those were hypothesis-driven, one locus at a time. What Voight and colleagues wanted was a map — a genome-wide, first-generation map of evolution still in progress.

The resource that made it possible was the International HapMap Project, which provided roughly eight hundred thousand single nucleotide polymorphisms — single-letter differences in the genetic code — assayed in two hundred nine unrelated individuals drawn from three continental populations: Yoruba from Nigeria, northern and western Europeans, and East Asians. That density of markers, spread across the whole genome, finally made a systematic search feasible. The method they developed to do this is called iHS — the integrated haplotype score — and it works by reading haplotype length as a proxy for time. A haplotype is a chunk of chromosome inherited together. The longer and more uniform it is, the less time recombination has had to chop it up. When a beneficial allele rises quickly, the haplotype it rides on stays long. The key observable is something called EHH, or extended haplotype homozygosity, which measures the probability that two chromosomes carrying the same allele are genetically identical across increasingly distant sites. EHH starts at one right at the focal variant and decays toward zero as you move outward.

Voight and colleagues integrated the area under that decay curve separately for the ancestral version of each variant and the derived version — the newer mutation. Then, they took the log ratio of those two integrated areas and standardized it by allele frequency. In plain terms, iHS is asking whether the derived allele sits on an unusually long haplotype compared to what you would expect for a variant at that frequency. Large negative values flag derived alleles with suspiciously long haplotypes — the signature of a sweep in progress. They computed iHS for over six hundred seventy thousand single nucleotide polymorphisms in Europeans, over seven hundred thousand in Yoruba, and over six hundred twenty thousand in East Asians, using a high-resolution recombination map built from linkage disequilibrium data. The sanity check is reassuring: the lactase region in Europeans, LCT, lights up brilliantly. That stretch spans about two thousand eight hundred kilobases, contains three hundred fifty-one out of five hundred ninety-four single nucleotide polymorphisms with an absolute iHS above two, and reaches a maximum iHS of four point nine — one of the most extreme values in the entire European dataset. That's a variant the field already knew was under strong recent selection, associated with the ability to digest dairy into adulthood, almost certainly driven by the spread of pastoralism. The method finds it cleanly.

Now scale that up to the whole genome. What Voight and colleagues found is that signals of recent positive selection are widespread across all three populations — but are mostly local. They scanned in one hundred kilobase windows, flagging windows in the top one percent of their empirical distribution as candidate sweep regions. Most of those windows were population-specific. That's the headline: adaptation over this timescale has been regional, shaped by the particular environments and pressures each population encountered. However, there was a statistically clear excess of windows shared between two or even all three populations, more than chance would predict. This suggests some sweeps are responses to challenges that humans faced broadly. Then comes the result that flips an earlier assumption. Low-resolution genome scans had sometimes suggested that recent selection was relatively sparse in sub-Saharan Africans. Voight and colleagues find the opposite: by several measures, some of their strongest signals come from the Yoruba sample.

The Yoruba sweep regions are narrower — average genetic spans of zero point thirty-two centimorgans, versus zero point fifty-two centimorgans in East Asians and Europeans — which likely explains why earlier coarser analyses missed them. Narrower regions mean older or stronger sweeps, not weaker ones. The age estimates, derived from haplotype spans assuming a twenty-five-year generation time, put average Yoruba sweeps at roughly ten thousand eight hundred years old versus about six thousand six hundred years in non-Africans. These events are recent by evolutionary standards, falling mainly within the Holocene — the last ten thousand years, the era of agriculture, changing diets, new pathogens, and dramatic shifts in population density and geography. The longest individual haplotypes point to the most extreme events. Near the GBA gene in East Asians, a selected haplotype extends one point thirty-nine centimorgans. Near NKX2-2 in Europeans, it extends one point twenty-five centimorgans. On chromosome five p fifteen in Yoruba, it extends zero point ninety-seven centimorgans in what the authors describe as a gene desert — the genetic target there remains unknown.

This brings up the functional biology. Voight and colleagues cataloged what kinds of genes sit near the strongest selection signals, and several themes emerge clearly. Metabolism is prominent: LCT for lactase in Europeans, the SI gene for sucrose processing in East Asians, MAN2A1 in Yoruba and East Asians, and multiple genes tied to fat metabolism including LEPR — the leptin receptor, involved in regulating body fat mass — in East Asians, and NCOA1 in Yoruba. Detoxification shows up strongly too, particularly in Europeans, where cytochrome P450 genes — enzymes that break down a wide range of compounds — are significantly enriched. This includes CYP3A5, CYP2E1, and CYP1A2, as well as four genes in a CYP cluster on chromosome one. Pigmentation genes appear in Europeans: OCA2, MYO5A, DTNBP1, and TYRP1. Neurological candidates include two microcephaly-associated genes, CDK5RAP2 in Yoruba and CENPJ in Europeans and East Asians, plus the serotonin transporter SLC6A4 in Europeans and East Asians, and SNTG1 across all three populations.

The genome-wide Gene Ontology analysis turned up enrichment in chemosensory perception, with p-values of zero point zero zero zero six in Europeans and zero point zero zero zero four in Yoruba. There is olfaction at similar significance, carbohydrate metabolism in East Asians at a p-value of zero point zero zero two, electron transport in Europeans at a p-value of zero point zero zero two, and reproduction-related categories with p-values around zero point zero zero three to zero point zero zero four. An unexpected finding in Yoruba is that the X chromosome is strikingly enriched, with fifteen genes mapped to top one percent windows in the Yoruba sample versus only six in Europeans and two in East Asians. Voight and colleagues are careful to state the central caveat plainly: for most of these signals, the actual phenotype under selection is unknown. Even when a candidate gene is identifiable, what it was selected for often remains unclear. These are starting points, not finished explanations. That caveat coexists with a forward-looking argument. If a variant has been rising under strong selection — simulations put the selection coefficients in the range of zero point zero one to zero point zero four — it must be producing real differences in fitness, which almost certainly means real differences in phenotype. And real phenotypic differences are exactly what genetic mapping studies need.

A few of the identified regions already have established medical relevance: CYP3A5 and salt-sensitive hypertension, the alcohol dehydrogenase cluster ADH and alcoholism susceptibility, and the chromosomal inversion on seventeen q twenty-one and fertility. On that basis, Voight and colleagues built a practical tool: genome-wide iHS scores for HapMap single nucleotide polymorphisms and a set of tag SNPs designed to capture the top approximately two hundred fifty selection signals in each population. The logic is that wherever selection has recently pushed an allele toward high frequency, that region is worth examining for associations with complex traits. The map that Voight, Kudaravalli, Wen, and Pritchard produced is a first-generation chart of where the human genome has been actively shaped in the last several thousand years. Most signals are local. Some are shared. The strongest, by some measures, come from Africa. The genes involved touch metabolism, detoxification, sensory processing, reproduction, and brain development. The timing puts most events squarely in the agricultural era. What drove each sweep, in most cases, remains to be determined. But the map exists now — and it shows, unmistakably, that human evolution did not stop. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Biochemistry, Genetics and Molecular Biology