Gene-Wide Analysis Detects Two New Susceptibility Genes for Alzheimer's Disease

Valentina Escott‐Price, Céline Bellenguez, Li‐San Wang, Seung‐Hoan Choi, Denise Harold, Lesley Jones, Peter Holmans, Amy Gerrish, Alexey Vedernikov, Alexander Richards, Anita L. DeStefano, Jean‐Charles Lambert, Carla A. Ibrahim‐Verbaas, Adam C. Naj, Rebecca Sims, Gyungah Jun, Joshua C. Bis, Gary W. Beecham, Benjamin Grenier‐Boley, Giancarlo Russo, Tricia A. Thornton‐Wells, Nicola Denning, Albert V. Smith, Vincent Chouraki, Charlene Thomas, M. Arfan Ikram, Diana Zélénika, Badri N. Vardarajan, Yoichiro Kamatani, Chiao‐Feng Lin, Helena Schmidt, Brian W. Kunkle, Melanie Dunstan, Maria Vronskaya, Andrew D. Johnson, Agustı́n Ruiz, Marie‐Thérèse Bihoreau, Christiane Reitz, Florence Pasquier, Paul Hollingworth, Olivier Hanon, Annette L. Fitzpatrick, Joseph D. Buxbaum, Dominique Campion, Paul K. Crane, Clinton T. Baldwin, Tim Becker, Vilmundur Guðnason, Carlos Cruchaga, David Craig, Najaf Amin, Claudine Berr, Oscar L. Lopez, Philip L. De Jager, Vincent Deramecourt, Janet Johnston, Denis A. Evans, Simon Lovestone, Luc Letenneur, Isabel Hernández, David C. Rubinsztein, Gudny Eiriksdottir, Kristel Sleegers, Alison Goate, Nathalie Fiévet, Matthew J. Huentelman, Michael Gill, Kristelle Brown, M. Ilyas Kamboh, Lina Keller, Pascale Barberger‐Gateau, Bernadette McGuinness, Eric B. Larson, Amanda Myers, Carole Dufouil, Stephen Todd, David Wallon, Seth Love, Ekaterina Rogaeva, John Gallacher, Peter St George‐Hyslop, Jordi Clarimón, Alberto Lleó, Anthony Bayer, Debby W. Tsuang, Lei Yu, Magda Tsolaki, Paola Bossù, Gianfranco Spalletta, Petra Proitsi, John Collinge, Sandro Sorbi, Florentino Sánchez-García, Nick C. Fox, John Hardy, María Cándida Déniz Naranjo, Paolo Bosco, Robert Clarke, Carol Brayne, Daniela GalimbertiView original
OverviewBalancedhelen voice
Picture a genome-wide association study the way it's usually run. A researcher scans millions of individual letters in the DNA — one position at a time, one letter at a time — comparing people with Alzheimer's to healthy controls, looking for the single spot where the two groups differ most. It's painstaking, and for two decades it was the only way. Then Escott-Price and colleagues asked a different question: what if you stopped reading letter by letter and started reading word by word — testing entire genes as a unit? That shift in the level of analysis, applied to the largest Alzheimer's genetics dataset ever assembled, is what this paper is built on. The gap that motivated the work is real and it's large. Late-onset Alzheimer's disease has a heritability estimated between fifty-six and seventy-nine percent, meaning genetics accounts for more than half of a person's risk. Yet standard genome-wide association studies had turned up only about twenty common susceptibility loci beyond the long-established APOE gene. That contrast — massive heritability, relatively few mapped variants — is what geneticists call the missing heritability problem. Part of the reason is structural. If a gene harbors several modest risk variants rather than one large one, no single position in the DNA will cross the strict significance threshold on its own. The signal is real, but it's scattered across the gene like sound spread through multiple quiet speakers. Single SNP tests — SNP stands for single-nucleotide polymorphism, a one-letter difference in the genetic code — simply aren't designed to add those whispers together. The International Genomics of Alzheimer's Project, known as IGAP, gave Escott-Price and colleagues the scale to do something about that. After imputation and quality filtering, Stage 1 of the analysis retained more than seven million autosomal SNPs drawn from seventeen thousand and eight Alzheimer's cases and thirty-seven thousand and one hundred fifty-four controls, pooled across four consortia. A second replication stage added eight thousand five hundred seventy-two cases and eleven thousand three hundred twelve controls — bringing the combined total to twenty-five thousand five hundred eighty cases and forty-eight thousand four hundred sixty-six controls. That's not a slight improvement on what came before. That's the kind of statistical power that lets you see faint signals. The gene-wide test the team used is an LD-adjusted Fisher method. LD stands for linkage disequilibrium — nearby variants tend to be inherited together, which means their p-values are correlated, not independent. The standard Fisher approach to combining p-values assumes independence, so a naive application would overstate the evidence. The adjusted version rescales the test statistic based on the pairwise correlations between SNPs, using one thousand Genomes data as the reference. Verbally, what the formula does is this: take each SNP's p-value, compute its natural logarithm, sum those logarithms across all SNPs in the gene, multiply by negative two — and then correct for the number of effectively independent markers. The result is a single number that represents the gene's collective evidence for association. After mapping seven million fifty-five thousand eight hundred eighty-one SNPs to twenty-five thousand three hundred ten testable genes, the team set a gene-wide significance threshold of two point five times ten to the negative six to account for the number of tests. Stage 1 did exactly what a well-designed analysis should do first: it recovered what was already known. The gene-wide approach reproduced genome-wide significant associations at established loci including CR1, BIN1, CLU, PICALM, ABCA7, SORL1, and APOE, among others. That validation matters — it confirms the method is working before you trust what it finds next. And beyond the replications, something else showed up. The distribution of gene-level p-values contained more small values than chance would predict. Escott-Price and colleagues computed the effective number of independent SNP tests across the genome — on the order of three point five to three point seven million — and compared the observed count of significant signals to the binomial expectation. There was a measurable excess. More real associations were hiding below the genome-wide threshold than the standard single SNP approach had surfaced. To find them, the team carried all loci with a Stage 1 gene-wide p-value below ten to the negative four into Stage 2 replication — eight hundred eighty-seven genes across four hundred forty-four independent loci. Two of those genes came back with genome-wide significant combined evidence. Neither had been identified by prior single SNP genome-wide association studies. Both tell a story that connects to biology the Alzheimer's field already had reasons to care about. The first is TP53INP1, sitting on chromosome 8. Its combined gene-wide p-value across both stages was one point four times ten to the negative six, driven by three SNPs that individually reached p-values below ten to the negative four. Those three SNPs show partial independence from each other — pairwise correlations of roughly r-squared of zero point six to zero point six five — suggesting more than one association signal within the gene. TP53INP1 encodes a protein involved in autophagy-dependent cell death and pro-apoptotic signaling, operating by changing the phosphorylation state of the p53 protein. It also modulates cell-to-extracellular matrix adhesion. In plain terms: this gene helps regulate when and how cells clear damaged material and, if necessary, die. Those are processes central to the protein aggregation pathology of Alzheimer's disease. The functional evidence goes beyond the association statistics. In a brain tissue comparison, TP53INP1 showed higher expression in one hundred thirty-seven Alzheimer's cases than in one hundred seventy-six controls, with a t-test p-value of zero point zero one three. The BRAINEAC brain dataset reported a cis-eQTL — a variant that changes how much of the gene's protein is produced in the brain — at a p-value of six point eight times ten to the negative six, located about seven point six kilobases upstream of the gene. Methylation data from a CpG island near the transcription start site added further regulatory support. This is a gene where the statistical signal, the brain expression data, and the known biology all point in the same direction. The second novel gene is IGHV1-67 on chromosome 14, with a combined p-value of seven point nine times ten to the negative eight — a stronger statistical signal than TP53INP1. IGHV1-67 is annotated as a pseudogene within the immunoglobulin heavy-chain variable region, the part of the genome involved in generating antibody diversity through a process called V(D)J recombination. The association rests on just two SNPs, and those two SNPs are in tight linkage disequilibrium with each other — r-squared of about zero point ninety-two — so the signal is essentially one independent association rather than multiple. The team did not find IGHV1-67 in the brain eQTL databases they queried, so the direct functional evidence is thinner here than for TP53INP1. But the gene's location in the immunoglobulin heavy-chain region connects it to adaptive immune function, a pathway the Alzheimer's field has increasingly implicated in disease susceptibility. That's the methodological lesson embedded in these findings. Neither TP53INP1 nor IGHV1-67 had a single SNP that crossed the standard genome-wide significance threshold of five times ten to the negative eight on its own. The gene-wide approach surfaced them precisely because it was designed to add the whispers. Of the twenty-seven genes with Stage 1 p-values at or below ten to the negative four, nine — thirty-three percent — replicated in Stage 2. Across all eight hundred eighty-seven tested genes, one hundred twenty-four showed nominal replication. The authors don't oversell this. The p-values are modest. Effect sizes are not large. But the pattern is consistent with a disease shaped by many variants of moderate effect distributed across a genome, rather than a few dominant mutations. The two new genes add to a cluster of signals pointing toward three biological pathways. Energy metabolism is implicated by genes like NDUFS3 and MTCH2. Protein degradation connects through ZNF3, which interacts with the ubiquitin-proteasomal system via BAG3. And immune function is now reinforced by IGHV1-67 alongside other immunoglobulin-region genes. Escott-Price and colleagues argue that when multiple genes and multiple analytical approaches converge on the same pathway, that convergence is meaningful for drug development — it's biological validation, not just a single correlation. They are careful to note that these are susceptibility loci, not demonstrated causal mechanisms. Which specific variant drives the signal and how it produces disease biology remain open questions. What the paper demonstrates cleanly is that the unit of analysis matters. Scanning letters one at a time left real biology on the table. Reading in words — in genes — changed what was visible. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Picture a genome-wide association study the way it's usually run. A researcher scans millions of individual letters in the DNA — one position at a time, one letter at a time — comparing people with Alzheimer's to healthy controls, looking for the single spot where the two groups differ most. It's painstaking, and for two decades it was the only way. Then Escott-Price and colleagues asked a different question: what if you stopped reading letter by letter and started reading word by word — testing entire genes as a unit? That shift in the level of analysis, applied to the largest Alzheimer's genetics dataset ever assembled, is what this paper is built on. The gap that motivated the work is real and it's large. Late-onset Alzheimer's disease has a heritability estimated between fifty-six and seventy-nine percent, meaning genetics accounts for more than half of a person's risk. Yet standard genome-wide association studies had turned up only about twenty common susceptibility loci beyond the long-established APOE gene. That contrast — massive heritability, relatively few mapped variants — is what geneticists call the missing heritability problem. Part of the reason is structural. If a gene harbors several modest risk variants rather than one large one, no single position in the DNA will cross the strict significance threshold on its own.

The signal is real, but it's scattered across the gene like sound spread through multiple quiet speakers. Single SNP tests — SNP stands for single-nucleotide polymorphism, a one-letter difference in the genetic code — simply aren't designed to add those whispers together. The International Genomics of Alzheimer's Project, known as IGAP, gave Escott-Price and colleagues the scale to do something about that. After imputation and quality filtering, Stage 1 of the analysis retained more than seven million autosomal SNPs drawn from seventeen thousand and eight Alzheimer's cases and thirty-seven thousand and one hundred fifty-four controls, pooled across four consortia. A second replication stage added eight thousand five hundred seventy-two cases and eleven thousand three hundred twelve controls — bringing the combined total to twenty-five thousand five hundred eighty cases and forty-eight thousand four hundred sixty-six controls. That's not a slight improvement on what came before. That's the kind of statistical power that lets you see faint signals. The gene-wide test the team used is an LD-adjusted Fisher method. LD stands for linkage disequilibrium — nearby variants tend to be inherited together, which means their p-values are correlated, not independent. The standard Fisher approach to combining p-values assumes independence, so a naive application would overstate the evidence.

The adjusted version rescales the test statistic based on the pairwise correlations between SNPs, using one thousand Genomes data as the reference. Verbally, what the formula does is this: take each SNP's p-value, compute its natural logarithm, sum those logarithms across all SNPs in the gene, multiply by negative two — and then correct for the number of effectively independent markers. The result is a single number that represents the gene's collective evidence for association. After mapping seven million fifty-five thousand eight hundred eighty-one SNPs to twenty-five thousand three hundred ten testable genes, the team set a gene-wide significance threshold of two point five times ten to the negative six to account for the number of tests. Stage 1 did exactly what a well-designed analysis should do first: it recovered what was already known. The gene-wide approach reproduced genome-wide significant associations at established loci including CR1, BIN1, CLU, PICALM, ABCA7, SORL1, and APOE, among others. That validation matters — it confirms the method is working before you trust what it finds next.

And beyond the replications, something else showed up. The distribution of gene-level p-values contained more small values than chance would predict. Escott-Price and colleagues computed the effective number of independent SNP tests across the genome — on the order of three point five to three point seven million — and compared the observed count of significant signals to the binomial expectation. There was a measurable excess. More real associations were hiding below the genome-wide threshold than the standard single SNP approach had surfaced. To find them, the team carried all loci with a Stage 1 gene-wide p-value below ten to the negative four into Stage 2 replication — eight hundred eighty-seven genes across four hundred forty-four independent loci. Two of those genes came back with genome-wide significant combined evidence. Neither had been identified by prior single SNP genome-wide association studies. Both tell a story that connects to biology the Alzheimer's field already had reasons to care about. The first is TP53INP1, sitting on chromosome 8. Its combined gene-wide p-value across both stages was one point four times ten to the negative six, driven by three SNPs that individually reached p-values below ten to the negative four. Those three SNPs show partial independence from each other — pairwise correlations of roughly r-squared of zero point six to zero point six five — suggesting more than one association signal within the gene.

TP53INP1 encodes a protein involved in autophagy-dependent cell death and pro-apoptotic signaling, operating by changing the phosphorylation state of the p53 protein. It also modulates cell-to-extracellular matrix adhesion. In plain terms: this gene helps regulate when and how cells clear damaged material and, if necessary, die. Those are processes central to the protein aggregation pathology of Alzheimer's disease. The functional evidence goes beyond the association statistics. In a brain tissue comparison, TP53INP1 showed higher expression in one hundred thirty-seven Alzheimer's cases than in one hundred seventy-six controls, with a t-test p-value of zero point zero one three. The BRAINEAC brain dataset reported a cis-eQTL — a variant that changes how much of the gene's protein is produced in the brain — at a p-value of six point eight times ten to the negative six, located about seven point six kilobases upstream of the gene. Methylation data from a CpG island near the transcription start site added further regulatory support. This is a gene where the statistical signal, the brain expression data, and the known biology all point in the same direction.

The second novel gene is IGHV1-67 on chromosome 14, with a combined p-value of seven point nine times ten to the negative eight — a stronger statistical signal than TP53INP1. IGHV1-67 is annotated as a pseudogene within the immunoglobulin heavy-chain variable region, the part of the genome involved in generating antibody diversity through a process called V(D)J recombination. The association rests on just two SNPs, and those two SNPs are in tight linkage disequilibrium with each other — r-squared of about zero point ninety-two — so the signal is essentially one independent association rather than multiple. The team did not find IGHV1-67 in the brain eQTL databases they queried, so the direct functional evidence is thinner here than for TP53INP1. But the gene's location in the immunoglobulin heavy-chain region connects it to adaptive immune function, a pathway the Alzheimer's field has increasingly implicated in disease susceptibility. That's the methodological lesson embedded in these findings. Neither TP53INP1 nor IGHV1-67 had a single SNP that crossed the standard genome-wide significance threshold of five times ten to the negative eight on its own. The gene-wide approach surfaced them precisely because it was designed to add the whispers.

Of the twenty-seven genes with Stage 1 p-values at or below ten to the negative four, nine — thirty-three percent — replicated in Stage 2. Across all eight hundred eighty-seven tested genes, one hundred twenty-four showed nominal replication. The authors don't oversell this. The p-values are modest. Effect sizes are not large. But the pattern is consistent with a disease shaped by many variants of moderate effect distributed across a genome, rather than a few dominant mutations. The two new genes add to a cluster of signals pointing toward three biological pathways. Energy metabolism is implicated by genes like NDUFS3 and MTCH2. Protein degradation connects through ZNF3, which interacts with the ubiquitin-proteasomal system via BAG3. And immune function is now reinforced by IGHV1-67 alongside other immunoglobulin-region genes. Escott-Price and colleagues argue that when multiple genes and multiple analytical approaches converge on the same pathway, that convergence is meaningful for drug development — it's biological validation, not just a single correlation. They are careful to note that these are susceptibility loci, not demonstrated causal mechanisms. Which specific variant drives the signal and how it produces disease biology remain open questions. What the paper demonstrates cleanly is that the unit of analysis matters. Scanning letters one at a time left real biology on the table. Reading in words — in genes — changed what was visible.

This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Medicine