A Systematic Review of Re-Identification Attacks on Health Data
De-identification of health data is broken. This claim has circulated through medical journals, law reviews, and congressional testimony for years, backed by a string of dramatic demonstrations. A researcher matches a governor's hospital records to a voter list using only a birth date, a ZIP code, and gender. Netflix users are unmasked from their movie ratings. AOL search logs expose people by name. The evidence, repeated often enough, starts to feel conclusive. El Emam, Jonker, Arbuckle, and Malin decided to actually count—not the headlines but the data. What they found is messier and more consequential than either side of the debate wants to confront. Here is why the question matters in both directions. Privacy law in most jurisdictions permits health data to flow for research, public health, and health services analysis without patient consent, but only if the data have been properly de-identified. Institutional review boards routinely waive consent requirements on that same basis. If de-identification genuinely fails, the consequences are serious. Patients lose privacy, custodians face pressure to require consent for every secondary use, and the incentive to de-identify at all could evaporate—meaning more identifiable data gets shared, not less. The U.S.
Department of Health and Human Services reported 252 breaches at health information custodians between late September 2009 and the end of 2010, each involving more than 500 records, exposing over 7.8 million patients. That is the world that bad de-identification produces. But the other failure is equally real. If policymakers and the public wrongly believe de-identification is always reversible, they lock away data that drives epidemiology, informs public health decisions, and increasingly trains medical artificial intelligence. Getting the actual evidence right is not an academic exercise. To achieve this, El Emam and colleagues ran a systematic review—a formal method for aggregating published research into a single, bias-conscious summary. They searched IEEE Xplore, the ACM Digital Library, and PubMed using terms like "anonymization," "de-identification," and "re-identification," covering everything through October 2010. The initial pool was 1,522 records. After two-stage screening with an inter-rater reliability check—the agreement statistic, called kappa, came in at 0.85, which is strong—fourteen distinct attacks survived to full analysis. These were studies that went beyond theoretical risk: they actually re-identified real individuals, and in eleven of the fourteen cases, verified those matches.
For each of the fourteen studies, the authors recorded six things: whether health data were involved, who the attacker was, what country the attack occurred in, how many records were re-identified, whether re-identification was verified, and—this is the critical variable—whether the data had been de-identified according to existing standards before the attack began. That last question is the engine of the entire paper. The headline number is this: across all fourteen attacks, the mean proportion of records successfully re-identified was 0.26—about 26 percent—with a 95 percent confidence interval running from 4.6 percent to 47.8 percent. For attacks specifically on health data, the mean rose to 0.34—about 34 percent—with a confidence interval from zero to 74.4 percent. A chi-squared test confirmed the studies were too heterogeneous to pool with a simple average, so the team used a random-effects estimator developed by Laird and Mosteller. Under that model, each study is weighted not just by its own sample size but also by the amount of variation across studies—which was substantial. The wide confidence intervals are not a statistical technicality; they are the honest signal that this body of evidence cannot support a confident verdict.
And then there is the number that changes everything: 0.00013. That is 0.013 percent. That is the re-identification success rate from the single attack in the entire review that targeted health data de-identified according to existing standards—specifically the HIPAA Safe Harbor criteria, which require the removal of eighteen categories of identifying information including names, geographic data smaller than state level, most date fields, and phone numbers. Only two of the fourteen attacks were carried out on data that had gone through standards-based de-identification. Only one of those two involved health data. That one attack found a success rate of 0.013 percent. The gap between 34 percent and 0.013 percent is not noise. It is the difference between attacking properly de-identified data and attacking data that was never de-identified in the first place. The 26 percent headline, in other words, is mostly a measure of what happens when you attack data that was inadequately protected before the attacker showed up. That finding gets sharper when you look at who was doing the attacking and what they were attacking. Of the fourteen successful re-identification exercises, eleven were performed by researchers—not criminals or foreign intelligence services, but academics running controlled demonstrations. Many had access to auxiliary registers, ground-truth verification, and jurisdiction-specific public data that a typical external attacker would not have.
Ten of the fourteen attacks were conducted by U.S.-based investigators, which matters because the availability of public voter rolls, property records, and commercial data in the United States makes certain linkage attacks far easier than they would be elsewhere. The datasets targeted ranged from Internet search logs and movie-rating databases—the AOL and Netflix examples—to health claims data with exact birth dates and five-digit ZIP codes. These attacks were demonstrations of theoretical vulnerability. They were designed to succeed, and they did. That does not make them representative of real-world risk. There is also the publication bias problem, and El Emam and colleagues quantify it carefully. Two mechanisms inflate the published record: researchers who fail to re-identify records rarely write that up, and journals are less interested in null results. The team ran a failsafe sensitivity analysis—asking how many unpublished low-success studies would be needed to pull the pooled mean down to a negligible level. The answer was uncomfortably small. If each unpublished study had a success rate of 0.1, only 23 such studies would be needed to push the confidence interval's upper bound below the pooled mean of 0.26. The funnel plots reinforced the concern: small studies scattered widely, often showing high re-identification rates, while attacks on larger databases clustered near zero.
No large database in the review showed a high re-identification proportion. The evidence base is fragile, and the paper states this plainly. So where does this leave us? El Emam and colleagues are careful not to overcorrect. Their conclusion is not that de-identification works fine and everyone should relax. Their conclusion is that the current evidence is insufficient to draw firm conclusions either way. The studies that generate the alarming headline rates are small, heterogeneous, and—overwhelmingly—attacks on data that was never properly de-identified to begin with. Remove those, and the picture changes dramatically. But the evidence base for how standards-based de-identification performs at scale, under realistic adversarial conditions, against modern auxiliary data sources, is nearly nonexistent. One standards-based health attack in the entire published literature. That is the gap. The practical implications cut in two directions. For data custodians, the paper recommends continuing to apply current de-identification best practices while pairing them with legal safeguards—data sharing agreements that prohibit re-identification, and data use agreements with accountability provisions. De-identification alone is not the last line of defense; it is one layer in a system.
For researchers and policymakers, the call is for better evidence: large-scale, standardized evaluations using data that has actually been de-identified by the book, under conditions that reflect what a realistic attacker could plausibly access. The stakes for getting that evidence are high. De-identified health data is what makes population health research possible. It feeds health services analysis, public health surveillance, and increasingly the machine learning models that clinicians are being asked to trust. The question of whether de-identification actually works is not a niche privacy dispute. It is a question about how much useful knowledge we can extract from the health system without violating the people inside it. El Emam and colleagues did not answer that question—but they did something equally valuable. They showed precisely how little the existing evidence actually tells us, and why the next generation of research needs to do better. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- To disclose or not disclose, is no longer the question – effect of AI-disclosed brand voice on brand authenticity and attitude
- Deciphering Interactions in Moving Animal Groups
- Predicting survival from colorectal cancer histology slides using deep learning: A retrospective multicenter study
- The manifold costs of being a non-native English speaker in science
- Towards a digital body: The virtual arm illusion
- Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening