Estimating the Frequency of Lyme Disease Diagnoses, United States, 2010–2018
Lyme disease is the most reported vector-borne illness in the United States. However, the official surveillance count is widely understood to be a substantial underestimate. That gap matters.
It matters to clinicians deciding whom to test and treat, to public health planners allocating prevention resources, and to patients trying to understand their own risk. Kugeler and colleagues at the Centers for Disease Control and Prevention set out to measure not just infections in the community, but something more immediately practical: how many people are actually being diagnosed and treated each year?
The causative agent is Borrelia burgdorferi, a spiral-shaped bacterium transmitted to humans by Ixodes ticks. The infection can involve multiple organ systems, but it is treatable with antimicrobial drugs, and most people recover fully, especially with early and appropriate treatment. Despite nearly three decades of national surveillance, though, the true clinical burden has remained unclear.
Here is the core problem. Lyme disease is nationally notifiable. This means physicians and labs are supposed to report confirmed cases to health departments, which roll up to the Centers for Disease Control and Prevention.
That system has produced a consistent picture of where disease concentrates — heavily in the Northeast, mid-Atlantic, and upper Midwest — and who gets it. Roughly 35,000 cases are reported through surveillance each year. But passive reporting is known to substantially undercount.
Prior work had estimated the underreporting factor for Lyme disease at somewhere between three and twelve. A previous claims-based analysis covering 2005 to 2010 put annual clinician-diagnosed cases at roughly 329,000 — nearly ten times the surveillance figure. What was happening in the decade that followed?
To answer that, Kugeler and colleagues turned to the IBM Watson Health MarketScan Commercial Claims and Encounters database, which contains insurance billing records for more than 25 million privately insured U.S. residents under the age of 65. Insurance claims are not a perfect window into disease, but they capture something surveillance misses entirely: the full volume of clinical encounters where a doctor diagnosed and treated a patient, regardless of whether that case was ever reported to public health.
The case definition the team used was straightforward in concept, though careful in execution. A patient was counted if they had both a Lyme disease diagnosis code — either ICD-9-CM code 088.81 or ICD-10-CM code A69.2x — and an associated prescription for an appropriate antibiotic lasting more than seven days. Repeat events in the same person in subsequent years were excluded to approximate incident, or new, cases rather than ongoing care.
Across the nine-year study period from 2010 to 2018, MarketScan contained 199 million person-years of observation. Within that pool, 118,780 people met the case definition.
But getting from that number to a national estimate required two corrections, both of which are worth understanding. The first was straightforward: MarketScan only covers people under 65. Surveillance data showed that 80.3 percent of confirmed and probable Lyme disease cases fell in that age group, so the team multiplied their standardized count by the reciprocal of 0.803 — about 1.25 — to scale up to all ages. That adjustment added roughly 25 percent to the count.
The second correction is subtler and more consequential. Even when a clinician diagnoses and treats Lyme disease, they often do not record the specific Lyme disease billing code in the chart. Kugeler and colleagues found this in three separate chart reviews: in New York, 114 of 273 records that met surveillance definitions contained the ICD-9 code; in Maryland, 84 of 236; and in Minnesota, 91 of 163.
Combined, 289 of 672 records — exactly 43 percent — had the specific code. That means the other 57 percent of diagnosed patients would be invisible to a claims-based analysis that did not account for it. So the team divided their age-corrected count by 0.43, a multiplier of about 2.33.
They more than doubled the estimate to account for cases that simply were not coded as Lyme disease, even when Lyme disease was what the doctor thought and treated.
After standardization to the U.S. population by five-year age group and state, applying the age correction and the undercoding correction, the analysis produced an estimate of approximately 476,000 clinician-diagnosed and treated Lyme disease cases per year, with a 95 percent credible interval of 405,000 to 547,000.
Let that land for a moment. Surveillance captures around 35,000 cases a year. This estimate is roughly 476,000.
That is about thirteen and a half times larger than the reported figure — and it represents not infections in the community at large, but patients who actually saw a doctor, received a diagnosis, and were prescribed antibiotics. This is clinical burden, not theoretical burden.
Geographically, the pattern is familiar: 81 percent of diagnoses in MarketScan came from residents of 14 high-incidence states in the Northeast, mid-Atlantic, and upper Midwest, with another 8 percent from adjoining states. But the remaining 19 percent came from low-incidence states — notably higher than the roughly 5 percent seen in surveillance data. The authors flag this as a signal worth watching: it may reflect increasing geographic spread of the disease, or it may reflect some degree of overdiagnosis in regions where Lyme disease is less common and clinicians may be less calibrated to its presentation.
The trend line also matters. The prior analysis covering 2005 to 2010 estimated about 329,000 annual diagnoses. This study's estimate for 2010 to 2018 is 476,000.
That increase — nearly 45 percent — parallels the rise seen in surveillance data over the same period, which adds some confidence that the trend is real rather than artifactual. Lyme disease appears to be genuinely growing in clinical footprint.
Now, the limitations. The undercoding correction is the most important methodological assumption in this study, and it rests on 672 chart records from three states. That is not a large foundation for a multiplier that more than doubles the final estimate.
Different coding patterns in different states or health systems could shift the result substantially. The authors are transparent about this — but it is the number to hold lightly.
The insurance claims themselves also have a ceiling. MarketScan covers privately insured patients under 65. People without commercial insurance, or on Medicaid, or uninsured, are not included in this data.
Access to health care differs across income levels and geographies, and that means some diagnosed cases are still being missed — not by undercoding, but by not entering the billing system at all. The estimate is almost certainly conservative in that respect.
There is also an inherent ambiguity the study cannot resolve. The gap between 476,000 diagnosed cases and 35,000 reported cases reflects two distinct forces pulling in opposite directions: underreporting of real infections to surveillance, and potential overdiagnosis in clinical practice. Claims data capture every prescription written, including prescriptions for patients who may not actually have Lyme disease.
Kugeler and colleagues are explicit that their analysis cannot separate how much of the gap comes from missed real cases versus presumptive treatment of patients who do not have the infection. That ambiguity matters for interpreting the number.
What should change? The authors point toward better electronic medical and laboratory systems as the clearest path forward — systems that could capture diagnoses more reliably without depending on the specific ICD code being entered. They also call for improved awareness among clinicians and the public to support early, accurate diagnosis and appropriate treatment.
The infrastructure for counting, in other words, needs to catch up with the scale of the problem.
Hundreds of thousands of people are being diagnosed with Lyme disease every year in the United States. That is the finding. Getting the count right is not an abstract epidemiological exercise — it is the prerequisite for knowing whether prevention resources are adequate, whether treatment guidelines are reaching the right populations, and whether the disease is spreading into new territory.
Kugeler and colleagues have given us the clearest estimate yet of that burden. The next step is building the systems that do not require this kind of reconstruction to produce it.
This lecture was created by ennepō.
Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field.
Read when you can. Listen when you want to.
Related lectures
- New insights into mechanisms behind miscarriage
- Long-Distance Delivery of Bacterial Virulence Factors by Pseudomonas aeruginosa Outer Membrane Vesicles
- A Distinct Macrophage Population Mediates Metastatic Breast Cancer Cell Extravasation, Establishment and Growth
- Liver Fluke Induces Cholangiocarcinoma
- The presence of tumor associated macrophages in tumor stroma as a prognostic marker for breast cancer patients
- Water, Sanitation, Hygiene, and Soil-Transmitted Helminth Infection: A Systematic Review and Meta-Analysis