The FTO (fat mass and obesity associated) gene codes for a novel member of the non-heme dioxygenase superfamily

Luis Sánchez‐Pulido, Miguel A. Andrade‐NavarroView original
OverviewBalancedmaya voice
One single-letter change in the human genome — a single nucleotide polymorphism buried in the first intron of a gene called FTO — shifts obesity risk across millions of people. The effect shows up in children as young as seven, measured not in lean mass but in fat. The genetic signal is unmistakable. The biology behind it was, for a long time, completely blank. That is the situation Sánchez-Pulido and Andrade-Navarro walked into. The FTO protein had no recognizable structural domain, no annotated function, and no convincing homologs in any database. Standard sequence searches came back empty. Without a functional clue, experimentalists had nowhere to start — no obvious assay to run, no cellular process to implicate, and no molecular handle on why a variant in this gene makes you more likely to carry excess fat. The paper they wrote is the story of how computational biology cracked that problem open. The first thing they tried was the obvious thing. They took the human FTO protein sequence and ran it against standard repositories — the National Center for Biotechnology Information, Ensembl, and the Joint Genome Institute. Basic Local Alignment Search Tool searches returned nothing useful. Prior reports had already flagged the absence of detectable homologs, so the team treated that failure not as a dead end but as a diagnostic: FTO might be a highly divergent member of a known family, one too far removed from its relatives for direct sequence comparison to bridge the gap. So they changed tactics. They gathered the weak, borderline hits that Basic Local Alignment Search Tool did return — sequences with just enough similarity to be suspicious — and built a multiple sequence alignment using T-Coffee, then refined it by hand. From that alignment, they constructed a profile: a hidden Markov model, or HMM, built with HMMer, that captures the statistical pattern of which amino acid positions are conserved and which vary across the candidate family. This is the key conceptual shift. Instead of asking "does this sequence look like FTO?", the model asks "does this sequence fit the pattern shared across everything that might be related to FTO?" The profile is far more sensitive than a pairwise comparison because it encodes what varies as well as what doesn't. They ran that HMM against the UniRef50 and UniRef90 databases using HMMsearch. Still uncertain about fold and function, they escalated one more time to profile-to-profile comparison — submitting the FTO HMM to the HHpred server, which compares it against HMM profiles built from sequences of known three-dimensional structure. This is where the trail broke open. The FTO N-terminal region, spanning roughly amino acids 57 to 324, matched the human AlkB homolog hABH3 with an E-value of three point two times ten to the minus twenty-one. It matched the bacterial AlkB protein with an E-value of three point one times ten to the minus twelve. The estimated error rate on those matches was below three percent. The sequence identity between FTO and those proteins was only about seventeen percent — low enough that Basic Local Alignment Search Tool would never flag it, but the profile-to-profile match was unambiguous. What FTO turned out to be is a non-heme dioxygenase — specifically, a member of the iron(II) and two-oxoglutarate-dependent dioxygenase superfamily. To decode that label: these enzymes use an iron atom as a catalytic cofactor and consume a molecule called two-oxoglutarate as a co-substrate to perform oxidative chemistry on organic molecules. AlkB, the structural relative identified here, is best known for repairing alkylated — chemically damaged — DNA and RNA. The family also includes enzymes involved in collagen synthesis, hypoxia signaling, and histone demethylation. The unifying chemistry is iron-enabled oxidation. The evidence for placing FTO in this family is not just the fold-recognition match. In the paper's multiple sequence alignment, the team identified specific conserved residues in FTO that correspond to the amino acids responsible for coordinating iron and binding two-oxoglutarate in characterized family members. Those residues appear in human FTO and in its homologs across other species. That conservation is mechanistically meaningful: these positions are under selection because they form the active center of the enzyme. Sánchez-Pulido and Andrade-Navarro conclude from this that both two-oxoglutarate and iron should be essential for FTO function. The evolutionary distribution of FTO adds another layer. Homologs appear consistently across vertebrates — from fish to mammals — but are absent from insects, worms, and fungi. That pattern points to a vertebrate origin. Then the picture gets strange: close homologs also show up in two groups of photosynthetic protists, the green alga Ostreococcus and the diatoms Thalassiosira and Phaeodactylum. The most parsimonious explanation, they argue, is horizontal gene transfer — FTO moving from vertebrates into those lineages independently. It is a provocative claim, and it implies this enzyme confers enough selective advantage to be retained when acquired. The localization data pull the story further in a specific direction. Human FTO has a predicted molecular mass of about fifty kilodaltons, and analysis using PSORTII identifies a seventeen-amino-acid bipartite nuclear localization signal spanning positions two to eighteen. A sliding-window analysis of the N-terminal region finds a peak of seven basic — lysine and arginine — residues out of seventeen at position ten, a concentration characteristic of nuclear localization signals. WolfPSORT returned scores split between cytoplasmic and nuclear-cytoplasmic localization, but the predicted signal and the conserved lysine and arginine-rich region in fish and mammalian homologs support the idea that FTO can enter the nucleus. Notably, this N-terminal extension is absent in the algal and diatom homologs — the ones that arrived by horizontal transfer. Put that together: FTO is predicted to be a nuclear-localized iron- and two-oxoglutarate-dependent oxidase, related to proteins that repair or modify nucleic acids. That combination immediately raises a question the paper poses but cannot yet answer — if FTO is modifying something inside the nucleus, what is it modifying? DNA? RNA? A histone? The sequence divergence from known family members is too great to read off the substrate from the structure alone. This is where the metabolic sensor hypothesis enters. The co-substrate two-oxoglutarate is not an obscure biochemical — it is an intermediate in the citric acid cycle, the cell's central energy-processing hub. An enzyme that requires two-oxoglutarate to function is, in principle, sensitive to the metabolic state of the cell: when energy metabolism shifts, two-oxoglutarate levels shift, and so does the enzyme's activity. Sánchez-Pulido and Andrade-Navarro suggest that FTO could act as a metabolic sensor regulatory protein — one that links nutritional or energetic status to some downstream nuclear process, and whose disruption contributes to an obese phenotype. They are explicit that this is a hypothesis, not a demonstrated mechanism. The substrate is unknown. The downstream pathway is uncharacterized. But the value of the hypothesis is not that it's proven — it's that it's testable. Before this paper, FTO was a sequence with no functional handle whatsoever. Now experimentalists have a specific enzymatic identity to work with. Does FTO bind iron? Does it consume two-oxoglutarate in a catalytic reaction? Does disrupting those conserved residues — the ones the alignment marks as shared across the non-heme dioxygenase family — abolish its activity? Does it localize to the nucleus via that N-terminal signal, and what happens to fat mass in animals where you mutate it? Those are tractable experiments. They follow directly and logically from what the computational analysis identified: the enzyme family, the predicted cofactors, the conserved active-site residues, and the localization signal. Sánchez-Pulido and Andrade-Navarro built that map from the failure of standard tools, through escalating profile-based methods, to a fold-recognition match so statistically strong it overcame seventeen percent sequence identity. A gene that looked like nothing in every database turned out to be a nuclear dioxygenase with a plausible connection to the cell's metabolic state. That is what rigorous computational biology, applied to a real biological mystery, can actually do. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

One single-letter change in the human genome — a single nucleotide polymorphism buried in the first intron of a gene called FTO — shifts obesity risk across millions of people. The effect shows up in children as young as seven, measured not in lean mass but in fat. The genetic signal is unmistakable. The biology behind it was, for a long time, completely blank. That is the situation Sánchez-Pulido and Andrade-Navarro walked into. The FTO protein had no recognizable structural domain, no annotated function, and no convincing homologs in any database. Standard sequence searches came back empty. Without a functional clue, experimentalists had nowhere to start — no obvious assay to run, no cellular process to implicate, and no molecular handle on why a variant in this gene makes you more likely to carry excess fat. The paper they wrote is the story of how computational biology cracked that problem open. The first thing they tried was the obvious thing. They took the human FTO protein sequence and ran it against standard repositories — the National Center for Biotechnology Information, Ensembl, and the Joint Genome Institute. Basic Local Alignment Search Tool searches returned nothing useful. Prior reports had already flagged the absence of detectable homologs, so the team treated that failure not as a dead end but as a diagnostic: FTO might be a highly divergent member of a known family, one too far removed from its relatives for direct sequence comparison to bridge the gap.

So they changed tactics. They gathered the weak, borderline hits that Basic Local Alignment Search Tool did return — sequences with just enough similarity to be suspicious — and built a multiple sequence alignment using T-Coffee, then refined it by hand. From that alignment, they constructed a profile: a hidden Markov model, or HMM, built with HMMer, that captures the statistical pattern of which amino acid positions are conserved and which vary across the candidate family. This is the key conceptual shift. Instead of asking "does this sequence look like FTO?", the model asks "does this sequence fit the pattern shared across everything that might be related to FTO?" The profile is far more sensitive than a pairwise comparison because it encodes what varies as well as what doesn't. They ran that HMM against the UniRef50 and UniRef90 databases using HMMsearch. Still uncertain about fold and function, they escalated one more time to profile-to-profile comparison — submitting the FTO HMM to the HHpred server, which compares it against HMM profiles built from sequences of known three-dimensional structure. This is where the trail broke open. The FTO N-terminal region, spanning roughly amino acids 57 to 324, matched the human AlkB homolog hABH3 with an E-value of three point two times ten to the minus twenty-one. It matched the bacterial AlkB protein with an E-value of three point one times ten to the minus twelve. The estimated error rate on those matches was below three percent.

The sequence identity between FTO and those proteins was only about seventeen percent — low enough that Basic Local Alignment Search Tool would never flag it, but the profile-to-profile match was unambiguous. What FTO turned out to be is a non-heme dioxygenase — specifically, a member of the iron(II) and two-oxoglutarate-dependent dioxygenase superfamily. To decode that label: these enzymes use an iron atom as a catalytic cofactor and consume a molecule called two-oxoglutarate as a co-substrate to perform oxidative chemistry on organic molecules. AlkB, the structural relative identified here, is best known for repairing alkylated — chemically damaged — DNA and RNA. The family also includes enzymes involved in collagen synthesis, hypoxia signaling, and histone demethylation. The unifying chemistry is iron-enabled oxidation. The evidence for placing FTO in this family is not just the fold-recognition match. In the paper's multiple sequence alignment, the team identified specific conserved residues in FTO that correspond to the amino acids responsible for coordinating iron and binding two-oxoglutarate in characterized family members. Those residues appear in human FTO and in its homologs across other species. That conservation is mechanistically meaningful: these positions are under selection because they form the active center of the enzyme. Sánchez-Pulido and Andrade-Navarro conclude from this that both two-oxoglutarate and iron should be essential for FTO function.

The evolutionary distribution of FTO adds another layer. Homologs appear consistently across vertebrates — from fish to mammals — but are absent from insects, worms, and fungi. That pattern points to a vertebrate origin. Then the picture gets strange: close homologs also show up in two groups of photosynthetic protists, the green alga Ostreococcus and the diatoms Thalassiosira and Phaeodactylum. The most parsimonious explanation, they argue, is horizontal gene transfer — FTO moving from vertebrates into those lineages independently. It is a provocative claim, and it implies this enzyme confers enough selective advantage to be retained when acquired. The localization data pull the story further in a specific direction. Human FTO has a predicted molecular mass of about fifty kilodaltons, and analysis using PSORTII identifies a seventeen-amino-acid bipartite nuclear localization signal spanning positions two to eighteen. A sliding-window analysis of the N-terminal region finds a peak of seven basic — lysine and arginine — residues out of seventeen at position ten, a concentration characteristic of nuclear localization signals.

WolfPSORT returned scores split between cytoplasmic and nuclear-cytoplasmic localization, but the predicted signal and the conserved lysine and arginine-rich region in fish and mammalian homologs support the idea that FTO can enter the nucleus. Notably, this N-terminal extension is absent in the algal and diatom homologs — the ones that arrived by horizontal transfer. Put that together: FTO is predicted to be a nuclear-localized iron- and two-oxoglutarate-dependent oxidase, related to proteins that repair or modify nucleic acids. That combination immediately raises a question the paper poses but cannot yet answer — if FTO is modifying something inside the nucleus, what is it modifying? DNA? RNA? A histone? The sequence divergence from known family members is too great to read off the substrate from the structure alone. This is where the metabolic sensor hypothesis enters. The co-substrate two-oxoglutarate is not an obscure biochemical — it is an intermediate in the citric acid cycle, the cell's central energy-processing hub. An enzyme that requires two-oxoglutarate to function is, in principle, sensitive to the metabolic state of the cell: when energy metabolism shifts, two-oxoglutarate levels shift, and so does the enzyme's activity.

Sánchez-Pulido and Andrade-Navarro suggest that FTO could act as a metabolic sensor regulatory protein — one that links nutritional or energetic status to some downstream nuclear process, and whose disruption contributes to an obese phenotype. They are explicit that this is a hypothesis, not a demonstrated mechanism. The substrate is unknown. The downstream pathway is uncharacterized. But the value of the hypothesis is not that it's proven — it's that it's testable. Before this paper, FTO was a sequence with no functional handle whatsoever. Now experimentalists have a specific enzymatic identity to work with. Does FTO bind iron? Does it consume two-oxoglutarate in a catalytic reaction? Does disrupting those conserved residues — the ones the alignment marks as shared across the non-heme dioxygenase family — abolish its activity? Does it localize to the nucleus via that N-terminal signal, and what happens to fat mass in animals where you mutate it? Those are tractable experiments. They follow directly and logically from what the computational analysis identified: the enzyme family, the predicted cofactors, the conserved active-site residues, and the localization signal. Sánchez-Pulido and Andrade-Navarro built that map from the failure of standard tools, through escalating profile-based methods, to a fold-recognition match so statistically strong it overcame seventeen percent sequence identity.

A gene that looked like nothing in every database turned out to be a nuclear dioxygenase with a plausible connection to the cell's metabolic state. That is what rigorous computational biology, applied to a real biological mystery, can actually do. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Chemistry