Diversity and scaleGenetic architecture of 2068 traits in the VA Million Veteran Program

Anurag Verma, Jennifer E. Huffman, Alex A Rodriguez, Mitchell Conery, Molei Liu, Yuk‐Lam Ho, Youngdae Kim, David Heise, Lindsay Guare, Vidul Ayakulangara Panickan, Helene Garcon, Franciel Linares, Lauren Costa, Ian Goethert, Ryan Tipton, Jacqueline Honerlaw, Laura Davies, Stacey B. Whitbourne, Jérémy Cohen, Daniel Posner, Rahul Sangar, Michael Murray, Xuan Wang, Daniel Dochtermann, Poornima Devineni, Yunling Shi, Tarak Nandi, Themistocles L. Assimes, Charles A. Brunette, Robert J. Carroll, Royce E. Clifford, Scott L. DuVall, Joel Gelernter, Adriana M. Hung, Sudha K. Iyengar, Jacob Joseph, Rachel L. Kember, Henry R. Kranzler, Colleen Morse Kripke, Daniel F. Levey, Shiuh‐Wen Luoh, Victoria C. Merritt, Cassie Overstreet, Joseph D. Deak, Struan F.A. Grant, Renato Polimanti, Panos Roussos, Gabrielle Shakt, Yan V. Sun, Noah L. Tsao, Sanan Venkatesh, Georgios Voloudakis, Amy C. Justice, Edmon Begoli, Rachel Ramoni, Georgia D. Tourassi, Saiju Pyarajan, Philip S. Tsao, Christopher J. O’Donnell, Sumitra Muralidhar, Jennifer Moser, Juan P. Casas, Alexander G. Bick, Wei Zhou, Tianxi Cai, Benjamin F. Voight, Kelly Cho, J. Michael Gaziano, Ravi Madduri, Scott M. Damrauer, Katherine P. LiaoView original
OverviewBalancedadam voice
If most of what we know about human genetics comes from people of European descent, who make up roughly 16 percent of the world's population, then our maps of disease risk are incomplete for the other 84 percent. If those maps are incomplete, the variants we think cause disease may not be the right ones. And if we have been pointing our fine-mapping tools at the wrong targets, the drugs and interventions built on that science could fail the people who need them most. A research team led by Anurag Verma and colleagues set out to fix this, one veteran at a time, across 635,969 people and 2,068 traits. The scale of the problem is not subtle. Ninety-five percent of participants in published genome-wide association studies, or GWASs, which are the systematic scans that link genetic variants to traits, are genetically similar to European reference populations. That's not a niche critique. It means the genetic risk maps guiding drug development, polygenic risk scores, and clinical research were drawn almost entirely from one slice of humanity. Verma and colleagues argue that the problem must be addressed at scale: you need large, well-powered samples across diverse ancestry groups both to discover associations that only appear in certain populations and to test whether the genetic architecture underlying common traits is actually shared or subtly different across groups. The vehicle for that effort is the VA Million Veteran Program, established in 2011. For this analysis, the authors used 635,969 participants, partitioned into four groups by genetic similarity to reference populations: 449,042 European, which is abbreviated to EUR, 121,177 African, which is abbreviated to AFR, 59,048 Admixed American, which is abbreviated to AMR, and 6,702 East Asian, which is abbreviated to EAS. The cohort is predominantly male, with only 8.8 percent female participants, and has a mean age of 61.9 years. Veterans bring something unusual to genetic research: deep, longitudinal electronic health records from the VA system, which allowed the team to define 1,854 binary traits and 214 quantitative traits drawn from diagnosis codes, lab values, vital signs, and enrollment questionnaires. That amounts to 2,068 distinct traits tested across populations, with more than 44 million genetic variants in play after imputation and quality control. Running 4,045 independent genome-wide association studies across all those traits and populations meant computing more than 350 billion variant-trait associations. An unmodified version of the standard statistical tool SAIGE would have required roughly 251 compute years to finish that job. The team implemented a graphics-processing-unit-optimized version of SAIGE and ran it on the Department of Energy's Oak Ridge Leadership Computing Facility, cutting the computational burden by a factor of 160 — completing the work in 14,286 GPU hours. That's fourteen days of wall time, not centuries. Now for what they found. Across all traits and populations, the multi-population meta-analysis identified 13,672 genomic risk loci, which are regions of the genome flagged as associated with a trait, for 1,270 traits, at a study-wide significance threshold of a p-value less than 4.6 times ten to the negative eleven. That threshold was calibrated to the number of statistically independent traits in the study, approximately 1,038. Here is the number that matters most: 1,608 of those 13,672 loci were only significant after including individuals from non-European populations. They would have been invisible in a European-only study. That's not a rounding error — it's more than ten percent of all discovered loci, gone, if the researchers had done what most genetic studies do and drawn only from EUR participants. The mechanism is concrete. Over half the variants analyzed in the meta-analysis were absent from the EUR-only genome-wide association studies, and roughly a quarter of all analyzed variants, which amounts to ten million of them, were present only in AFR participants. Allele frequency differences drive specific, clinically meaningful discoveries. A rare intronic variant near the long noncoding RNA gene PCAT2, called rs72725854, has a minor allele frequency of about 6 percent in AFR and just 0.06 percent in EUR. It's the strongest prostate cancer signal in the study. A locus associated with keloid scarring, which are the raised, overgrown scars that form at wound sites, reached significance in the AFR group at a p-value of 2.2 times ten to the negative eleven. Keloid scarring is three times more prevalent in AFR than in EUR participants in this cohort. The variant responsible, rs76024540, is located in and around the SLC22A18 gene and its antisense partner SLC22A18AS, and has a minor allele frequency of about 11 percent in AFR and is monomorphic, meaning it is absent entirely, in EUR. You literally cannot find it if you don't look in the right population. Finding a locus is step one. Knowing which specific variant within that locus is actually causing the effect is a harder problem, and that's where fine-mapping comes in. Fine-mapping uses patterns of linkage disequilibrium, which is the tendency of nearby variants to be inherited together, to narrow a flagged region down to the most likely causal variant or variants. Verma and colleagues applied the Sum of Single Effects model, known as SuSiE, using exact in-sample linkage disequilibrium matrices, and required a posterior inclusion probability above 0.95 to call a variant high-confidence. The result: 6,318 distinct causal variants were fine-mapped across 613 traits. One-third of those, which amounts to 2,069 variant-trait associations, were identified specifically in non-European participants. And diversity doesn't just add discovery here; it adds precision. When the same locus was fine-mapped in both AFR and EUR participants, the AFR credible sets, which are the sets of candidate causal variants, were significantly smaller. A Wilcoxon signed-rank test comparing credible set sizes gave a p-value of 2.26 times ten to the negative ten in favor of AFR being more precise. After downsampling EUR participants to match AFR sample composition, the AFR precision advantage held at a p-value of 3.8 times ten to the negative fifty-two. The reason for this is that African populations tend to have smaller linkage disequilibrium blocks, which are shorter stretches where variants travel together, so the correlation structure acts like a finer grid, pinpointing the causal site more accurately. The rs76024540 keloid variant illustrates this perfectly. It sits in an Activity-by-Contact enhancer connecting to the promoters of SLC22A18 and SLC22A18AS in multiple cell types, including skin. One variant, one locus, identified because the right population was included. What does the broader genetic architecture look like across these groups? Mostly shared, with critical exceptions at the edges. Verma and colleagues used a tool called Popcorn to compare genetic effects across populations for the same traits. Of 236 traits that were heritable in both AFR and EUR, 168 showed significant cross-population genetic correlation. Height had the strongest correlation between AFR and EUR, with a genetic correlation coefficient of 0.66. Type 2 diabetes came in at 0.65. However, some traits diverged sharply: skin cancer showed a genetic correlation of just 0.05 between populations, and anemia of chronic disease was 0.08. Those low numbers aren't statistical noise; they reflect genuine biological differences in how variants map to traits across groups. Pleiotropy, which is when one variant or gene affects multiple traits, was pervasive throughout the atlas. The team nominated 15,596 trait-gene associations and found 2,279 genes each linked to two or more genetically independent traits, totaling 6,711 pleiotropic associations. The APOE gene alone connected to 29 traits. Using a Poisson regression framework to model how often genes accumulated trait associations, the relationship between the number of Gene Ontology terms per gene and its count of associated traits was highly significant at a p-value of 1.4 times ten to the negative seventeen. When they looked for genes that were pleiotropic outliers in AFR versus EUR, which are genes with substantially more trait nominations in AFR than their EUR count would predict, the leaders were APOL1, HBB, and CD36, all genes with well-known biological relevance to conditions more prevalent in populations of African ancestry. Verma and colleagues are direct about the limitations. The veteran cohort is older and predominantly male. The EAS group, with just 6,702 participants, remains underpowered for many traits. Even in this large study, EUR participants still outnumber the others combined, meaning differential statistical power persists. These are real constraints on what can be concluded within each subgroup. However, the 1,608 loci discovered only because diverse populations were included constitute a proof of concept that is hard to argue with. Every genome-wide association study that excludes non-European participants is not just failing on equity grounds — it is leaving specific, findable, clinically relevant answers undetected. The atlas that Verma and colleagues built, spanning 2,068 traits and 635,969 people, demonstrates what those answers look like when you actually go looking for them. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

If most of what we know about human genetics comes from people of European descent, who make up roughly 16 percent of the world's population, then our maps of disease risk are incomplete for the other 84 percent. If those maps are incomplete, the variants we think cause disease may not be the right ones. And if we have been pointing our fine-mapping tools at the wrong targets, the drugs and interventions built on that science could fail the people who need them most. A research team led by Anurag Verma and colleagues set out to fix this, one veteran at a time, across 635,969 people and 2,068 traits. The scale of the problem is not subtle. Ninety-five percent of participants in published genome-wide association studies, or GWASs, which are the systematic scans that link genetic variants to traits, are genetically similar to European reference populations. That's not a niche critique. It means the genetic risk maps guiding drug development, polygenic risk scores, and clinical research were drawn almost entirely from one slice of humanity. Verma and colleagues argue that the problem must be addressed at scale: you need large, well-powered samples across diverse ancestry groups both to discover associations that only appear in certain populations and to test whether the genetic architecture underlying common traits is actually shared or subtly different across groups.

The vehicle for that effort is the VA Million Veteran Program, established in 2011. For this analysis, the authors used 635,969 participants, partitioned into four groups by genetic similarity to reference populations: 449,042 European, which is abbreviated to EUR, 121,177 African, which is abbreviated to AFR, 59,048 Admixed American, which is abbreviated to AMR, and 6,702 East Asian, which is abbreviated to EAS. The cohort is predominantly male, with only 8.8 percent female participants, and has a mean age of 61.9 years. Veterans bring something unusual to genetic research: deep, longitudinal electronic health records from the VA system, which allowed the team to define 1,854 binary traits and 214 quantitative traits drawn from diagnosis codes, lab values, vital signs, and enrollment questionnaires. That amounts to 2,068 distinct traits tested across populations, with more than 44 million genetic variants in play after imputation and quality control.

Running 4,045 independent genome-wide association studies across all those traits and populations meant computing more than 350 billion variant-trait associations. An unmodified version of the standard statistical tool SAIGE would have required roughly 251 compute years to finish that job. The team implemented a graphics-processing-unit-optimized version of SAIGE and ran it on the Department of Energy's Oak Ridge Leadership Computing Facility, cutting the computational burden by a factor of 160 — completing the work in 14,286 GPU hours. That's fourteen days of wall time, not centuries. Now for what they found. Across all traits and populations, the multi-population meta-analysis identified 13,672 genomic risk loci, which are regions of the genome flagged as associated with a trait, for 1,270 traits, at a study-wide significance threshold of a p-value less than 4.6 times ten to the negative eleven. That threshold was calibrated to the number of statistically independent traits in the study, approximately 1,038. Here is the number that matters most: 1,608 of those 13,672 loci were only significant after including individuals from non-European populations. They would have been invisible in a European-only study. That's not a rounding error — it's more than ten percent of all discovered loci, gone, if the researchers had done what most genetic studies do and drawn only from EUR participants.

The mechanism is concrete. Over half the variants analyzed in the meta-analysis were absent from the EUR-only genome-wide association studies, and roughly a quarter of all analyzed variants, which amounts to ten million of them, were present only in AFR participants. Allele frequency differences drive specific, clinically meaningful discoveries. A rare intronic variant near the long noncoding RNA gene PCAT2, called rs72725854, has a minor allele frequency of about 6 percent in AFR and just 0.06 percent in EUR. It's the strongest prostate cancer signal in the study. A locus associated with keloid scarring, which are the raised, overgrown scars that form at wound sites, reached significance in the AFR group at a p-value of 2.2 times ten to the negative eleven. Keloid scarring is three times more prevalent in AFR than in EUR participants in this cohort. The variant responsible, rs76024540, is located in and around the SLC22A18 gene and its antisense partner SLC22A18AS, and has a minor allele frequency of about 11 percent in AFR and is monomorphic, meaning it is absent entirely, in EUR. You literally cannot find it if you don't look in the right population.

Finding a locus is step one. Knowing which specific variant within that locus is actually causing the effect is a harder problem, and that's where fine-mapping comes in. Fine-mapping uses patterns of linkage disequilibrium, which is the tendency of nearby variants to be inherited together, to narrow a flagged region down to the most likely causal variant or variants. Verma and colleagues applied the Sum of Single Effects model, known as SuSiE, using exact in-sample linkage disequilibrium matrices, and required a posterior inclusion probability above 0.95 to call a variant high-confidence. The result: 6,318 distinct causal variants were fine-mapped across 613 traits. One-third of those, which amounts to 2,069 variant-trait associations, were identified specifically in non-European participants. And diversity doesn't just add discovery here; it adds precision. When the same locus was fine-mapped in both AFR and EUR participants, the AFR credible sets, which are the sets of candidate causal variants, were significantly smaller. A Wilcoxon signed-rank test comparing credible set sizes gave a p-value of 2.26 times ten to the negative ten in favor of AFR being more precise. After downsampling EUR participants to match AFR sample composition, the AFR precision advantage held at a p-value of 3.8 times ten to the negative fifty-two.

The reason for this is that African populations tend to have smaller linkage disequilibrium blocks, which are shorter stretches where variants travel together, so the correlation structure acts like a finer grid, pinpointing the causal site more accurately. The rs76024540 keloid variant illustrates this perfectly. It sits in an Activity-by-Contact enhancer connecting to the promoters of SLC22A18 and SLC22A18AS in multiple cell types, including skin. One variant, one locus, identified because the right population was included. What does the broader genetic architecture look like across these groups? Mostly shared, with critical exceptions at the edges. Verma and colleagues used a tool called Popcorn to compare genetic effects across populations for the same traits. Of 236 traits that were heritable in both AFR and EUR, 168 showed significant cross-population genetic correlation. Height had the strongest correlation between AFR and EUR, with a genetic correlation coefficient of 0.66. Type 2 diabetes came in at 0.65. However, some traits diverged sharply: skin cancer showed a genetic correlation of just 0.05 between populations, and anemia of chronic disease was 0.08. Those low numbers aren't statistical noise; they reflect genuine biological differences in how variants map to traits across groups.

Pleiotropy, which is when one variant or gene affects multiple traits, was pervasive throughout the atlas. The team nominated 15,596 trait-gene associations and found 2,279 genes each linked to two or more genetically independent traits, totaling 6,711 pleiotropic associations. The APOE gene alone connected to 29 traits. Using a Poisson regression framework to model how often genes accumulated trait associations, the relationship between the number of Gene Ontology terms per gene and its count of associated traits was highly significant at a p-value of 1.4 times ten to the negative seventeen. When they looked for genes that were pleiotropic outliers in AFR versus EUR, which are genes with substantially more trait nominations in AFR than their EUR count would predict, the leaders were APOL1, HBB, and CD36, all genes with well-known biological relevance to conditions more prevalent in populations of African ancestry. Verma and colleagues are direct about the limitations. The veteran cohort is older and predominantly male. The EAS group, with just 6,702 participants, remains underpowered for many traits. Even in this large study, EUR participants still outnumber the others combined, meaning differential statistical power persists. These are real constraints on what can be concluded within each subgroup.

However, the 1,608 loci discovered only because diverse populations were included constitute a proof of concept that is hard to argue with. Every genome-wide association study that excludes non-European participants is not just failing on equity grounds — it is leaving specific, findable, clinically relevant answers undetected. The atlas that Verma and colleagues built, spanning 2,068 traits and 635,969 people, demonstrates what those answers look like when you actually go looking for them. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Biochemistry, Genetics and Molecular Biology