Biases in the Experimental Annotations of Protein Function and Their Effect on Our Understanding of Protein Function Space
A curator sits down to annotate a protein. Not one protein, but twenty-five thousand of them, all traceable to a single published paper. That ratio is not a thought experiment; it is a real feature of how the world's protein databases are built. Schnoes and colleagues spent a paper asking what it means for everything built on top of those databases. Here is the system: proteins get their functional labels from curators who read the scientific literature and place findings into a structured vocabulary called the Gene Ontology, or GO. The GO organizes biological knowledge into a hierarchy — cellular components, biological processes, molecular functions. Curators attach terms from that hierarchy to proteins, along with an evidence code recording how the association was made. When a paper is the source, the annotation record includes that paper's identifier. The result is a compilation called UniProt-GOA, which is the primary experimental record of what proteins do.
Schnoes and colleagues analyzed the December 2011 release of UniProt-GOA and found that the distribution of proteins annotated per article is exponential — a heavy-tailed curve with a long tail reaching far to the right. The team fit a line to the log-log plot of that distribution and got a p-value below one times ten to the negative eighteen, with an R-squared of about 0.72. The practical upshot is stark: zero point one four percent of annotating articles — one hundred and eight papers out of seventy-six thousand — provide the source of annotations for twenty-five percent of all experimentally annotated proteins in the database. One hundred and eight papers. Twenty-five percent of the proteins. The question is what those papers are doing, and the answer shapes everything that follows. The dominant studies are high-throughput experiments: mass spectrometry pipelines, microscopy-based imaging assays, and genome-wide RNA interference screens. When the team analyzed the top fifty annotating articles in detail, they found that twenty-seven of them are essentially mass-spectrometry studies using cellular fractionation to localize proteins to compartments. The most common evidence label assigned to those papers corresponds to "protein separation followed by fragment identification evidence" — a mass-spect fingerprint.
Imaging assays come second, and RNA interference loss-of-function screens come third. These are powerful technologies, but each of them is built to find a specific kind of information. Mass spectrometry tells you where a protein lives in the cell, and RNA interference screens tell you whether knocking out a gene disrupts a developmental process. Neither was designed to tell you what a protein actually does at the molecular level. That limitation shows up directly in the distribution of GO terms those experiments produce. High-throughput articles assign fifty-seven percent of their GO terms to Cellular Component — where is this protein in the cell — and thirty-eight percent to Biological Process, with only five percent going to Molecular Function — what this protein actually does biochemically. Low- and medium-throughput studies look completely different: roughly twenty-two to twenty-six percent Molecular Function, fifty-one to fifty-seven percent Biological Process, and only seventeen to twenty-five percent Cellular Component. The high-throughput experiments are not just different in scale; they are different in kind. Two information-theoretic metrics make this concrete. The first is edge count: how many steps from the root of the GO hierarchy to the term assigned. A shallow term like "catalytic activity" is one edge from the root.
A specific term like "haloalkane dehalogenase activity" is five. The deeper the term, the more diagnostic. The second metric is information content — the logarithm of the inverse of a term's frequency in the corpus. Rarer terms carry more information because they distinguish proteins from a smaller background. Both metrics tell the same story. Across Molecular Function and Biological Process, high-throughput articles produce significantly shallower, less informative annotations than low-throughput ones. In Cellular Component, the edge-count trend disappears — the terms are already shallow and stay shallow — but the information-content metric still shows a significant drop in the high-throughput cohort. More papers, more proteins, less learned about each one. There is also a redundancy problem. Within the top fifty annotating articles, large fractions of the proteins appear more than once — annotated by multiple high-throughput studies that drew on shared experimental resources like common RNA interference libraries. Caenorhabditis elegans shows sixty percent redundancy in that set, Arabidopsis thaliana forty-seven percent, and Mus musculus forty-six percent. The same proteins, annotated again and again, by experiments using the same tools, finding the same things.
So far, this is a story about what experimental databases do and don't contain. But the real stakes are larger because most protein annotations are not experimental; they are computational. Algorithms infer function from sequence similarity, transferring labels from experimentally characterized proteins to uncharacterized ones. That transfer process means experimental biases don't stay in the experimental layer; they propagate. A database skewed toward subcellular location and developmental phenotypes becomes a training set skewed the same way. And algorithms trained on biased training sets learn biased priors. Schnoes and colleagues point directly to this in the context of computational function prediction. Most predictors rely on prior probabilities derived from available annotations. If those annotations over-represent shallow, high-throughput terms, the algorithms will be pulled toward predicting shallow, high-throughput terms — even for proteins where deeper functional information might in principle be inferred. Their paper notes that during the 2011 Critical Assessment of Function Annotation challenge, roughly twenty percent of proteins annotated under Molecular Function were labeled "protein binding," a shallow, nearly uninformative term whose prominence came largely from high-throughput assays. Benchmarks built from such data reward algorithms for getting the easy, common terms right. The signal of progress can be partly illusory.
And it is worth being clear about which organisms are most exposed. The fraction of experimentally annotated proteins that come exclusively from high-throughput studies varies dramatically by species. For Schizosaccharomyces pombe, it is about sixty-two percent. For Caenorhabditis elegans, forty-seven percent. For Homo sapiens, around thirty-five percent. By contrast, Saccharomyces cerevisiae sits at eight point six percent and Escherichia coli K-12 at just five point two percent. The well-studied model organisms with decades of low-throughput biochemistry behind them are relatively protected. The others are almost entirely dependent on what a handful of mass-spectrometry and RNA interference papers happened to find. What can be done? The authors are honest that some of this is unavoidable. Every experimental assay is limited by what it can detect. A mass-spectrometry pipeline will always produce localization data, not mechanistic detail. The technology shapes the knowledge. But they point to one concrete tool: the Evidence Code Ontology, or ECO, which goes considerably beyond the twenty standard GO evidence codes.
ECO supplies specific labels for the actual assay, such as "microscopy," "RNA interference," and "protein separation followed by fragment identification," rather than the broad brush of "inferred by direct assay." When annotations carry that level of detail, downstream users can filter, weight, or flag high-throughput annotations differently. An algorithm building a training set could, in principle, treat a mass-spectrometry localization annotation differently from a directly assayed enzymatic activity. A benchmark designer could check whether their positive set is dominated by a single study's output. The paper also calls for better communication among curators, computational biologists, and experimentalists. It calls for annotation records to include the number of proteins annotated by the source paper, so the throughput level is immediately visible to anyone who pulls the record. These are not grand calls for more funding or more data. They are specific, implementable changes to how information is recorded and transmitted. The final point Schnoes and colleagues leave the reader with is epistemological. Knowing a bias exists is not the same as fixing it, but it is the necessary first step to not being fooled by it. A protein annotated exclusively because a mass-spectrometry study found it in a particular cellular fraction tells us far less about its molecular function than a protein characterized by years of biochemical work.
Those two proteins can sit next to each other in a database, carrying the same evidence code, looking equally well understood. They are not. The difference is in the provenance — and right now, most of the infrastructure for using protein function data makes that difference invisible. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Integrated analysis of ultra-deep proteomes in cortex, cerebrospinal fluid and serum reveals a mitochondrial signature in Alzheimer’s disease
- Generation of Persister Cells of Pseudomonas aeruginosa and Staphylococcus aureus by Chemical Treatment and Evaluation of Their Susceptibility to Membrane-Targeting Agents
- Chemical control of structure and guest uptake by a conformationally mobile porous material
- Gene expression changes in mononuclear cells in patients with metabolic syndrome after acute intake of phenol-rich virgin olive oil
- Heavy Metal Contaminations in Herbal Medicines: Determination, Comprehensive Risk Assessments, and Solutions
- Importance of c-Type cytochromes for U(VI) reduction by Geobacter sulfurreducens