Modeling 3D Facial Shape from DNA

Peter Claes, Denise K. Liberton, Katleen Daniels, Kerri Matthes Rosana, Ellen E. Quillen, Laurel N. Pearson, Brian McEvoy, Marc Bauchet, Arslan A. Zaidi, Wei Yao, Hua Tang, Gregory S. Barsh, Devin Absher, David A. Puts, Jorge Rocha, Sandra Beleza, Rinaldo Wellerson Pereira, Gareth Baynam, Paul Suetens, Dirk Vandermeulen, Jennifer K. Wagner, James S. Boster, Mark D. ShriverView original
OverviewBalancededdie_stirling voice
If you handed a forensic investigator a DNA sample from an unknown person — no name, no photo, no witness — could they sketch the face? Not just a rough demographic profile, but a three-dimensional prediction of cheekbone height, nose width, and the specific curve of a jaw. Hold that question. Claes and colleagues built a method that takes a serious step toward making this possible. What they found along the way reveals something profound about how genes, ancestry, and sex are inscribed into the geometry of a human face. Human faces are extraordinarily diverse, and for decades, genetics struggled to study that diversity systematically. Most genome-wide association studies reduced faces to a sparse set of manually placed anatomical landmarks — about a dozen fixed points — and extracted simple distances and angles from them. The problem is that this drastically oversimplifies facial shape. A face isn't a handful of measurements; it's a continuous surface. The difference between two people's noses, or the subtle protrusion of a brow ridge, exists in the geometry between those landmarks. Principal component analysis, or PCA, can describe face-space more richly, but it produces dozens of statistically independent axes of variation. Linking each of those axes to genes separately creates a noisy, nearly uninterpretable mess. The phenotype — the face — was too high-dimensional for the standard genetic toolset, and that mismatch stalled the field. Claes and colleagues tackled both sides of the problem simultaneously. Instead of placing landmarks by hand, they covered each face with seven thousand one hundred fifty quasi-landmarks — automated, spatially dense vertices that tile the entire facial surface. Each face becomes twenty-one thousand four hundred fifty numbers: X, Y, and Z coordinates for every vertex. They also removed asymmetry by averaging each face with its mirror image, so what remained was the pure bilateral shape. Working with five hundred ninety-two participants from three cohorts — the United States, Brazil, and Cape Verde — all with mixed West African and European ancestry, they ran PCA on those coordinates and found that forty-four principal components captured ninety-eight percent of the variation in face shape, with a reconstruction error of just zero point two millimeters per quasi-landmark on average. Forty-four components is still too many to link cleanly to genes. That's where their key innovation comes in: BRIM, bootstrapped response-based imputation modeling. The concept is elegant. Instead of asking which of the forty-four components is associated with ancestry, BRIM asks a different question — what single axis through face-space best predicts ancestry? It compresses the entire multivariate facial response into one scalar summary called a response-based imputed predictor, or RIP variable. RIP-A for ancestry, RIP-S for sex, and RIP-G for a specific gene. BRIM uses a partial least squares regression framework, with nested leave-one-out cross-validation to prevent overfitting and bootstrapping to monitor stability across iterations. Crucially, because you can model sex and ancestry out before computing a genetic RIP, the genetic effects you find are not confounded by the large correlated effects of someone's ancestry. The results for ancestry and sex are striking. RIP-A, the ancestry axis, correlates with genomic ancestry measured from a sixty-eight marker ancestry-informative panel at an r value of zero point eight one — meaning roughly two-thirds of the variation in RIP-A is explained by genomic ancestry alone. And it manifests in specific places on the face. RIP-A primarily shapes the nose and lips. On the European end of the axis, there’s increased surface area at the sides of the nose and at the front of the chin. On the West African end, there's greater surface area for the nostrils and lips, with a highly concave curvature at the top of the philtrum. The nasal bridge, supraorbital ridges, and chin all show measurable curvature differences. At its peak location, RIP-A explains just over forty percent of the variance in local facial shape. RIP-S, the sex axis, focuses on different regions: the supraorbital ridges, nasal bridge, zygomatics, and cheeks. The area under the curve for classifying self-reported sex was zero point nine nine four. At the total-face level, sex independently explains twelve point nine percent of facial variation, while genomic ancestry explains nine point six percent independently. These are substantial effects for a biological trait this complex. But here's the part that really resonates. Claes and colleagues didn't just show that their algorithm detected these axes; they demonstrated that human observers detect the same ones. Participants viewed animated three-dimensional faces, stripped of pigmentation and hair so only geometry remained, and rated proportional West African ancestry on a zero-to-hundred scale and femininity on a seven-point scale. The correlation between those human ratings and RIP-A was zero point eight five four, while for RIP-S and femininity ratings, it was zero point eight six zero. Both correlations were significant well beyond a p-value below zero point zero zero zero one. The faces encode ancestry and sex-related geometry in ways that both an algorithm and an untrained human eye can perceive. The signal isn't mathematical noise; it's socially legible shape. Then comes the genetics. Claes and colleagues tested seventy-six ancestry-informative single nucleotide polymorphisms — or SNPs — located in forty-six craniofacial candidate genes. For each SNP, they computed a RIP-G: the facial axis most associated with that variant, after removing sex and ancestry. Twenty-four of the seventy-six RIP-G variables, representing twenty different genes, showed associations at a p-value below zero point one. Given the strong prior evidence for these candidate genes and the small expected effects of single variants on a continuous trait, the team deemed this threshold appropriate. The gene-level results are specific. The SNP rs1074265 in the SLC35D1 gene produced a RIP-G with its largest effects around the eyes and periorbital region, with a maximum local R-squared of eleven point sixty-eight percent. The SNP rs13267109 in the FGFR1 gene — part of the fibroblast growth factor receptor family, involved in skeletal and craniofacial development — showed broader effects on the forehead, supraorbital ridges, nasal bridge, and mouth corners, with a maximum R-squared of fifteen point sixteen percent. The SNP rs2724626 in the LRP6 gene produced a clear change in lip shape — vermilion prominence and thickness — with a maximum R-squared of ten point ten percent. Multiple SNPs in the DNMT3B and SATB2 genes exhibited similar localized effects, and genes in the WNT and FGF signaling pathways produced overlapping signatures in adjacent facial regions. These aren't associations with disease or dysmorphology; these are normal faces, normal variation, healthy people — and specific variants in specific genes are nudging specific facial geometries in measurable, visualizable directions. The team is careful about what this proves and what it doesn't. The current models are statistical tendencies at the population level, not individual predictions. The study populations were admixed West African and European cohorts, and the paper emphasizes that many more populations must be studied before anyone can claim the results generalize. The candidate gene set was modest and purposeful — forty-six genes selected for prior evidence, not a genome-wide sweep. The method is a proof of concept, a framework, not a finished forensic tool. Claes and colleagues assert directly that much more work is needed before anyone can know how many genes will be required to estimate the shape of a face in any useful way. What the work does establish is that the framework functions. BRIM transforms an intractably high-dimensional phenotype into a series of targeted, biologically interpretable questions. It allows researchers to ask, cleanly and separately: what does sex do to this face? What does ancestry do? What does this particular allele do, with sex and ancestry held constant? For each question, the answer is a specific, visualizable axis through face-space — not an abstract statistical result, but a shape you could render and observe. The potential extensions the team mentions include forensic facial approximation from crime-scene DNA, diagnostic tools for craniofacial syndromes, and basic science into the developmental genetics of the face. The face turns out to be deeply legible — to algorithms, to human observers, and now to genetics. Claes and colleagues provided researchers a way to read that script, one gene at a time. The question that opened this research — can a DNA sample sketch a face? — doesn't have a full answer yet. But for the first time, it has a method. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

If you handed a forensic investigator a DNA sample from an unknown person — no name, no photo, no witness — could they sketch the face? Not just a rough demographic profile, but a three-dimensional prediction of cheekbone height, nose width, and the specific curve of a jaw. Hold that question. Claes and colleagues built a method that takes a serious step toward making this possible. What they found along the way reveals something profound about how genes, ancestry, and sex are inscribed into the geometry of a human face. Human faces are extraordinarily diverse, and for decades, genetics struggled to study that diversity systematically. Most genome-wide association studies reduced faces to a sparse set of manually placed anatomical landmarks — about a dozen fixed points — and extracted simple distances and angles from them. The problem is that this drastically oversimplifies facial shape. A face isn't a handful of measurements; it's a continuous surface. The difference between two people's noses, or the subtle protrusion of a brow ridge, exists in the geometry between those landmarks. Principal component analysis, or PCA, can describe face-space more richly, but it produces dozens of statistically independent axes of variation. Linking each of those axes to genes separately creates a noisy, nearly uninterpretable mess. The phenotype — the face — was too high-dimensional for the standard genetic toolset, and that mismatch stalled the field.

Claes and colleagues tackled both sides of the problem simultaneously. Instead of placing landmarks by hand, they covered each face with seven thousand one hundred fifty quasi-landmarks — automated, spatially dense vertices that tile the entire facial surface. Each face becomes twenty-one thousand four hundred fifty numbers: X, Y, and Z coordinates for every vertex. They also removed asymmetry by averaging each face with its mirror image, so what remained was the pure bilateral shape. Working with five hundred ninety-two participants from three cohorts — the United States, Brazil, and Cape Verde — all with mixed West African and European ancestry, they ran PCA on those coordinates and found that forty-four principal components captured ninety-eight percent of the variation in face shape, with a reconstruction error of just zero point two millimeters per quasi-landmark on average. Forty-four components is still too many to link cleanly to genes. That's where their key innovation comes in: BRIM, bootstrapped response-based imputation modeling. The concept is elegant. Instead of asking which of the forty-four components is associated with ancestry, BRIM asks a different question — what single axis through face-space best predicts ancestry? It compresses the entire multivariate facial response into one scalar summary called a response-based imputed predictor, or RIP variable. RIP-A for ancestry, RIP-S for sex, and RIP-G for a specific gene.

BRIM uses a partial least squares regression framework, with nested leave-one-out cross-validation to prevent overfitting and bootstrapping to monitor stability across iterations. Crucially, because you can model sex and ancestry out before computing a genetic RIP, the genetic effects you find are not confounded by the large correlated effects of someone's ancestry. The results for ancestry and sex are striking. RIP-A, the ancestry axis, correlates with genomic ancestry measured from a sixty-eight marker ancestry-informative panel at an r value of zero point eight one — meaning roughly two-thirds of the variation in RIP-A is explained by genomic ancestry alone. And it manifests in specific places on the face. RIP-A primarily shapes the nose and lips. On the European end of the axis, there’s increased surface area at the sides of the nose and at the front of the chin. On the West African end, there's greater surface area for the nostrils and lips, with a highly concave curvature at the top of the philtrum. The nasal bridge, supraorbital ridges, and chin all show measurable curvature differences. At its peak location, RIP-A explains just over forty percent of the variance in local facial shape.

RIP-S, the sex axis, focuses on different regions: the supraorbital ridges, nasal bridge, zygomatics, and cheeks. The area under the curve for classifying self-reported sex was zero point nine nine four. At the total-face level, sex independently explains twelve point nine percent of facial variation, while genomic ancestry explains nine point six percent independently. These are substantial effects for a biological trait this complex. But here's the part that really resonates. Claes and colleagues didn't just show that their algorithm detected these axes; they demonstrated that human observers detect the same ones. Participants viewed animated three-dimensional faces, stripped of pigmentation and hair so only geometry remained, and rated proportional West African ancestry on a zero-to-hundred scale and femininity on a seven-point scale. The correlation between those human ratings and RIP-A was zero point eight five four, while for RIP-S and femininity ratings, it was zero point eight six zero. Both correlations were significant well beyond a p-value below zero point zero zero zero one. The faces encode ancestry and sex-related geometry in ways that both an algorithm and an untrained human eye can perceive. The signal isn't mathematical noise; it's socially legible shape.

Then comes the genetics. Claes and colleagues tested seventy-six ancestry-informative single nucleotide polymorphisms — or SNPs — located in forty-six craniofacial candidate genes. For each SNP, they computed a RIP-G: the facial axis most associated with that variant, after removing sex and ancestry. Twenty-four of the seventy-six RIP-G variables, representing twenty different genes, showed associations at a p-value below zero point one. Given the strong prior evidence for these candidate genes and the small expected effects of single variants on a continuous trait, the team deemed this threshold appropriate. The gene-level results are specific. The SNP rs1074265 in the SLC35D1 gene produced a RIP-G with its largest effects around the eyes and periorbital region, with a maximum local R-squared of eleven point sixty-eight percent. The SNP rs13267109 in the FGFR1 gene — part of the fibroblast growth factor receptor family, involved in skeletal and craniofacial development — showed broader effects on the forehead, supraorbital ridges, nasal bridge, and mouth corners, with a maximum R-squared of fifteen point sixteen percent.

The SNP rs2724626 in the LRP6 gene produced a clear change in lip shape — vermilion prominence and thickness — with a maximum R-squared of ten point ten percent. Multiple SNPs in the DNMT3B and SATB2 genes exhibited similar localized effects, and genes in the WNT and FGF signaling pathways produced overlapping signatures in adjacent facial regions. These aren't associations with disease or dysmorphology; these are normal faces, normal variation, healthy people — and specific variants in specific genes are nudging specific facial geometries in measurable, visualizable directions. The team is careful about what this proves and what it doesn't. The current models are statistical tendencies at the population level, not individual predictions. The study populations were admixed West African and European cohorts, and the paper emphasizes that many more populations must be studied before anyone can claim the results generalize. The candidate gene set was modest and purposeful — forty-six genes selected for prior evidence, not a genome-wide sweep. The method is a proof of concept, a framework, not a finished forensic tool. Claes and colleagues assert directly that much more work is needed before anyone can know how many genes will be required to estimate the shape of a face in any useful way.

What the work does establish is that the framework functions. BRIM transforms an intractably high-dimensional phenotype into a series of targeted, biologically interpretable questions. It allows researchers to ask, cleanly and separately: what does sex do to this face? What does ancestry do? What does this particular allele do, with sex and ancestry held constant? For each question, the answer is a specific, visualizable axis through face-space — not an abstract statistical result, but a shape you could render and observe. The potential extensions the team mentions include forensic facial approximation from crime-scene DNA, diagnostic tools for craniofacial syndromes, and basic science into the developmental genetics of the face. The face turns out to be deeply legible — to algorithms, to human observers, and now to genetics. Claes and colleagues provided researchers a way to read that script, one gene at a time. The question that opened this research — can a DNA sample sketch a face? — doesn't have a full answer yet. But for the first time, it has a method. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Mathematics