Solving the missing heritability problem

Alexander I. YoungView original
OverviewBalancedadam voice
For most of the twentieth century, scientists used twins to measure how much of a trait — height, disease risk, intelligence — was written in the genes. The method is elegant in its simplicity: if identical twins are more similar than non-identical twins, that excess similarity implies a genetic contribution. The answers kept coming back high. For height, twin studies in European samples gave heritability estimates between seventy-three and eighty-one percent. Then, around two thousand seven, genome-wide association studies arrived — technologies that could scan millions of actual DNA variants across thousands of people. The field expected confirmation. What it got was a shock. By two thousand ten, roughly forty genetic variants had been linked to height. Together, they explained about five percent of height variation. Five percent, against a twin-study promise of nearly eighty. That chasm — between what twins implied and what DNA delivered — became known as the missing heritability problem. Alexander Young's paper is a systematic account of where that missing heritability actually went. To understand the gap, you first have to understand what those early genome-wide association studies, or GWAS, were actually measuring. Genotyping arrays — the chips used in GWAS — assayed roughly two hundred fifty thousand common variants, called single nucleotide polymorphisms, or SNPs. These are positions in the genome where people routinely differ by a single DNA letter. The chips capture common variants well but miss rare ones almost entirely. The statistical bar for declaring a variant "significant" was set very high because you're running millions of simultaneous tests and need to guard against false positives. That threshold, Young explains, meant studies were powered only to find common variants with relatively strong effects. Here is the critical distinction that reframes everything: a variant not found in an early GWAS is not the same as a variant that doesn't exist. Many common variants with small effects on height were genuinely there, genuinely contributing — they just fell below the detection threshold because the sample sizes were too small. The GWAS wasn't wrong. It was looking at a real signal through a blurry lens. The methodological pivot that clarified this came from a technique called GREML — Genomic Relatedness Restricted Maximum Likelihood. The key insight behind GREML is to stop asking which variants are individually significant and start asking something more collective: how much of the similarity between two people's traits can be predicted from how similar their entire genomes are? You build a matrix of pairwise genome-wide genetic similarity across thousands of unrelated individuals, then relate that to pairwise trait similarity, and estimate how much variance the chip-measured SNPs collectively explain. Yang and colleagues applied this to height in two thousand ten. The result was striking. Where GWAS had found five percent, GREML found roughly forty-five percent — explained by the same common variants on the same chips. That number, forty-five percent, is what Young calls chip heritability or SNP heritability: the fraction of trait variance tagged by the genetic variation captured on a genotyping array. It's not the same as total heritability. However, it told the field something important — most of what twins were detecting was spread across huge numbers of common variants, each with tiny effects, most too small to clear a significance threshold individually. The next step was to bring in rarer variants. Imputation is a statistical technique that uses known patterns of variant correlation — called linkage disequilibrium, the tendency of nearby variants to be inherited together — to infer variants that weren't directly measured on the chip. When Yang and colleagues applied imputation in a two thousand fifteen analysis using GREML-LDMS, the heritability estimate for height rose from forty-five to fifty-six percent. More variants, more variance explained. But still a gap with the twin estimate. That gap pointed toward an even rarer class of variants — ones that imputation can't reliably recover because they're so rare they're not well correlated with anything on the chip. Whole-genome sequencing takes the next step by directly measuring every variant in a person's genome, including the very rare ones. Wainschtein and colleagues applied GREML to whole-genome sequence data — a version called GREML-WGS — and reported a height heritability estimate of zero point seventy-nine, with a standard error of zero point zero nine. That lands squarely in the range of twin-study estimates. On its face, it looks like the gap is closed. Young urges caution, though, and the reason matters. Getting a precise estimate from GREML-WGS requires enormous samples — roughly forty thousand individuals, Young suggests, to distinguish between competing values with confidence. There’s also a confound called population stratification: systematic differences in ancestry between individuals that can make non-causal variants look like they’re associated with a trait, simply because ancestry and trait mean happen to covary. Rare variants have especially complex spatial distributions, and sharing of very rare variants often implies recent common ancestry — and potentially shared environment — making stratification particularly hard to correct without family data. This is where family-based methods become essential. Young describes a technique called Sib-Regression, which exploits random variation in how much genetic material siblings share due to the shuffling of chromosomes during reproduction. Because siblings share similar environments but random genetic portions, within-family analyses remove confounds from population stratification, assortative mating — where people tend to partner with others of similar trait values — and indirect genetic effects, meaning the genetic influence of relatives on your environment rather than on you directly. A related method, relatedness disequilibrium regression, or RDR, generalizes this to all relative pairs, gaining statistical precision while keeping the same robustness. These methods give more conservative answers. For height, RDR in Iceland produced an estimate of fifty-five percent, with a standard error of about four point four percent. Combined Sib-Regression estimates came in around sixty-eight percent. Both are noticeably below the twin-study range of seventy-three to eighty-one percent, and below the GREML-WGS estimate of seventy-nine percent. That gap between family-based estimates and twin-based estimates suggests that twin studies themselves may be inflated — by genetic interactions, shared environments, or assortative mating that the classical twin model doesn't fully account for. There is a practical payoff to all of this beyond pure measurement. Polygenic scores — predictions of a person's trait value or disease risk derived from their genome-wide genetic data — depend on understanding which variants matter and by how much. Young reports that polygenic scores for educational attainment currently explain eleven to thirteen percent of variation in that trait. But within-family analyses reveal that at least half of that predictive signal comes from sources other than direct genetic effects: population stratification, assortative mating, and what Young calls indirect genetic effects of relatives, sometimes called genetic nurture. A score trained on population-level data picks up all of these together. Only within-family designs can separate genuine predictive signal from inherited confound. So where did the missing heritability go? Young's answer is that it was never in one place. The largest portion was always there in common variants — distributed across thousands of loci, each with effects too small for early GWAS to detect individually, but collectively substantial. A second portion sits in rare variants that chips couldn't tag and imputation couldn't reliably recover; whole-genome sequencing suggests these are real contributors, though the precise amount remains uncertain. A third portion reflects methodological limitations — twin estimates inflated by indirect genetic effects, assortative mating, and possibly non-additive genetic interactions that the classical twin model attributes to heritability but shouldn't. What remains genuinely open are contributions from non-additive genetic effects — cases where the effect of one variant depends on another — and from gene-environment interactions too complex for current population designs to resolve cleanly. These are not assumed to be large, but they haven't been ruled out. The trajectory of the field, though, is toward resolution. Each generation of method — twin studies, early GWAS, GREML on chips, GREML with imputation, GREML-WGS, family-based regression — illuminated a different slice of the same underlying architecture. The estimates have been climbing: from five percent with identified GWAS hits, to forty-five percent with chip GREML, to fifty-six percent with imputation, to around sixty-eight to seventy-nine percent with sequencing and family designs, depending on which approach you trust most. The range tells you something real about remaining uncertainty. But the direction is unambiguous. The heritability was there all along. The tools were just catching up to it. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

For most of the twentieth century, scientists used twins to measure how much of a trait — height, disease risk, intelligence — was written in the genes. The method is elegant in its simplicity: if identical twins are more similar than non-identical twins, that excess similarity implies a genetic contribution. The answers kept coming back high. For height, twin studies in European samples gave heritability estimates between seventy-three and eighty-one percent. Then, around two thousand seven, genome-wide association studies arrived — technologies that could scan millions of actual DNA variants across thousands of people. The field expected confirmation. What it got was a shock. By two thousand ten, roughly forty genetic variants had been linked to height. Together, they explained about five percent of height variation. Five percent, against a twin-study promise of nearly eighty. That chasm — between what twins implied and what DNA delivered — became known as the missing heritability problem. Alexander Young's paper is a systematic account of where that missing heritability actually went. To understand the gap, you first have to understand what those early genome-wide association studies, or GWAS, were actually measuring. Genotyping arrays — the chips used in GWAS — assayed roughly two hundred fifty thousand common variants, called single nucleotide polymorphisms, or SNPs. These are positions in the genome where people routinely differ by a single DNA letter.

The chips capture common variants well but miss rare ones almost entirely. The statistical bar for declaring a variant "significant" was set very high because you're running millions of simultaneous tests and need to guard against false positives. That threshold, Young explains, meant studies were powered only to find common variants with relatively strong effects. Here is the critical distinction that reframes everything: a variant not found in an early GWAS is not the same as a variant that doesn't exist. Many common variants with small effects on height were genuinely there, genuinely contributing — they just fell below the detection threshold because the sample sizes were too small. The GWAS wasn't wrong. It was looking at a real signal through a blurry lens. The methodological pivot that clarified this came from a technique called GREML — Genomic Relatedness Restricted Maximum Likelihood. The key insight behind GREML is to stop asking which variants are individually significant and start asking something more collective: how much of the similarity between two people's traits can be predicted from how similar their entire genomes are? You build a matrix of pairwise genome-wide genetic similarity across thousands of unrelated individuals, then relate that to pairwise trait similarity, and estimate how much variance the chip-measured SNPs collectively explain.

Yang and colleagues applied this to height in two thousand ten. The result was striking. Where GWAS had found five percent, GREML found roughly forty-five percent — explained by the same common variants on the same chips. That number, forty-five percent, is what Young calls chip heritability or SNP heritability: the fraction of trait variance tagged by the genetic variation captured on a genotyping array. It's not the same as total heritability. However, it told the field something important — most of what twins were detecting was spread across huge numbers of common variants, each with tiny effects, most too small to clear a significance threshold individually. The next step was to bring in rarer variants. Imputation is a statistical technique that uses known patterns of variant correlation — called linkage disequilibrium, the tendency of nearby variants to be inherited together — to infer variants that weren't directly measured on the chip. When Yang and colleagues applied imputation in a two thousand fifteen analysis using GREML-LDMS, the heritability estimate for height rose from forty-five to fifty-six percent. More variants, more variance explained. But still a gap with the twin estimate.

That gap pointed toward an even rarer class of variants — ones that imputation can't reliably recover because they're so rare they're not well correlated with anything on the chip. Whole-genome sequencing takes the next step by directly measuring every variant in a person's genome, including the very rare ones. Wainschtein and colleagues applied GREML to whole-genome sequence data — a version called GREML-WGS — and reported a height heritability estimate of zero point seventy-nine, with a standard error of zero point zero nine. That lands squarely in the range of twin-study estimates. On its face, it looks like the gap is closed. Young urges caution, though, and the reason matters. Getting a precise estimate from GREML-WGS requires enormous samples — roughly forty thousand individuals, Young suggests, to distinguish between competing values with confidence. There’s also a confound called population stratification: systematic differences in ancestry between individuals that can make non-causal variants look like they’re associated with a trait, simply because ancestry and trait mean happen to covary. Rare variants have especially complex spatial distributions, and sharing of very rare variants often implies recent common ancestry — and potentially shared environment — making stratification particularly hard to correct without family data.

This is where family-based methods become essential. Young describes a technique called Sib-Regression, which exploits random variation in how much genetic material siblings share due to the shuffling of chromosomes during reproduction. Because siblings share similar environments but random genetic portions, within-family analyses remove confounds from population stratification, assortative mating — where people tend to partner with others of similar trait values — and indirect genetic effects, meaning the genetic influence of relatives on your environment rather than on you directly. A related method, relatedness disequilibrium regression, or RDR, generalizes this to all relative pairs, gaining statistical precision while keeping the same robustness. These methods give more conservative answers. For height, RDR in Iceland produced an estimate of fifty-five percent, with a standard error of about four point four percent. Combined Sib-Regression estimates came in around sixty-eight percent. Both are noticeably below the twin-study range of seventy-three to eighty-one percent, and below the GREML-WGS estimate of seventy-nine percent. That gap between family-based estimates and twin-based estimates suggests that twin studies themselves may be inflated — by genetic interactions, shared environments, or assortative mating that the classical twin model doesn't fully account for.

There is a practical payoff to all of this beyond pure measurement. Polygenic scores — predictions of a person's trait value or disease risk derived from their genome-wide genetic data — depend on understanding which variants matter and by how much. Young reports that polygenic scores for educational attainment currently explain eleven to thirteen percent of variation in that trait. But within-family analyses reveal that at least half of that predictive signal comes from sources other than direct genetic effects: population stratification, assortative mating, and what Young calls indirect genetic effects of relatives, sometimes called genetic nurture. A score trained on population-level data picks up all of these together. Only within-family designs can separate genuine predictive signal from inherited confound. So where did the missing heritability go? Young's answer is that it was never in one place. The largest portion was always there in common variants — distributed across thousands of loci, each with effects too small for early GWAS to detect individually, but collectively substantial.

A second portion sits in rare variants that chips couldn't tag and imputation couldn't reliably recover; whole-genome sequencing suggests these are real contributors, though the precise amount remains uncertain. A third portion reflects methodological limitations — twin estimates inflated by indirect genetic effects, assortative mating, and possibly non-additive genetic interactions that the classical twin model attributes to heritability but shouldn't. What remains genuinely open are contributions from non-additive genetic effects — cases where the effect of one variant depends on another — and from gene-environment interactions too complex for current population designs to resolve cleanly. These are not assumed to be large, but they haven't been ruled out. The trajectory of the field, though, is toward resolution. Each generation of method — twin studies, early GWAS, GREML on chips, GREML with imputation, GREML-WGS, family-based regression — illuminated a different slice of the same underlying architecture. The estimates have been climbing: from five percent with identified GWAS hits, to forty-five percent with chip GREML, to fifty-six percent with imputation, to around sixty-eight to seventy-nine percent with sequencing and family designs, depending on which approach you trust most. The range tells you something real about remaining uncertainty. But the direction is unambiguous. The heritability was there all along. The tools were just catching up to it.

This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Mathematics