Social Contacts and Mixing Patterns Relevant to the Spread of Infectious Diseases
If you want to understand how a respiratory epidemic actually spreads, you have to know who meets whom, for how long, and where. For years, modelers filled in those blanks with educated guesses, using synthetic contact matrices calibrated after the fact to match case curves. Useful, but shaky.
Then Mossong and colleagues did something simple and radical. Instead of assuming social mixing, they measured it directly, the same way across eight European countries.
Picture this: One day, from five in the morning to the next five in the morning, on a randomly assigned day of the week. A paper diary in your hand. Every person you touch skin-to-skin or talk to in a real, two-way conversation of at least three words gets written down.
You record the age and sex of that person, where you were—home, work, school, leisure, transport—how long it lasted, and how often you usually meet. Parents did it for little kids; older kids did it themselves. One participant per household, from Belgium to Poland, in local languages, with the same basic rules.
By the end, seven thousand two hundred ninety people had recorded ninety-seven thousand nine hundred four contacts. That’s a lot of ink. On average, each person logged thirteen point four contacts in that day.
Countries differed in the headline number—Germany came in low, Italy high—but the shape of who met whom looked surprisingly similar. Before we get to the patterns, hold onto that scale and the standardization. It’s what lets the results travel.
The big picture is easy to see once you plot age against age. The contact matrix lights up along the diagonal. Teenagers and young adults tend to see other teenagers and young adults.
The five-to-twenty-four age group forms a thick band of same-age mixing, showing up in Belgium, Finland, Great Britain, and everywhere they looked. Two fainter diagonals sit off to the side. One is kids with adults roughly thirty to thirty-nine years older—home life, parents and their children.
The mirror of that is older kids with middle-aged adults. Those lines matter for transmission, but they’re significantly lighter than that main assortative stripe. Among adults, you see a flatter plateau, driven by many lower-intensity contacts at work.
Intensity tracks context and cadence. Long or frequent interactions are much more likely to be physical. Short or rare ones, not so much.
In the pooled data, roughly three-quarters of home contacts involved touch. In schools and leisure settings, it was about half. Travel and other public-facing spaces dropped to around a third.
If you zoom in on duration and "how often," the patterns snap into focus: about seventy-five percent of first-time contacts that lasted less than five minutes were nonphysical, while a large majority of multi-hour, daily relationships involved touch. You don’t need a PhD to guess that; what’s valuable here is the size and consistency of the measurement.
The team didn’t just eyeball it. They put the attributes into a formal pattern search—association rule mining, the same kind of tool that tells a supermarket that diapers and beer co-occur in shopping carts. Using a minimum support of half a percent of all contacts—about five hundred encounters in this dataset—a significance threshold of one percent, and rules up to three attributes long, they pulled out ninety-nine robust associations.
Many of the high-lift rules paired long duration with daily frequency and physical contact. On the flip side, the short, first-time exchanges clustered as nonphysical. The math doesn’t change common sense; it tests it at scale and tells you which exceptions are rare.
Now, raw counts are messy. People differ wildly: a handful of very gregarious folks, many with modest days, and some professions where contact numbers explode. The distribution is skewed, and in a few countries, respondents were even told to cap how many frequent professional contacts to record.
So when Mossong’s group modeled the number of contacts per person, they didn’t use a simple Poisson distribution, which expects variance to march in lockstep with the mean. They used a negative binomial distribution, which has a separate knob for overdispersion—the extra, real-world scatter. In their regression, that knob, the dispersion parameter alpha, landed at zero point three six with a confidence interval from zero point three four to zero point three seven, a clear sign that the Poisson would have been too tight.
They also had to handle the diaries that were right-censored at twenty-nine contacts, so the likelihood splits observations into uncensored and censored contributions and weights each diary according to how typical that person’s age and household size are in national census data. If you want the equation in words: they modeled the expected number of contacts as the exponential of a linear combination of covariates—age, sex, household size, country, weekday—and then accounted for the fact that some diaries only told you "twenty-nine or more." It’s the statistical equivalent of not pretending you saw more detail than you did.
Two quick details from that regression tell you why context matters. First, weekdays are busier than Sundays. Depending on the country, people reported roughly thirty to forty percent more contacts on a weekday than on Sunday.
Second, age and household size are not just background information; they move the needle, with children and larger households associated with more reported contacts. Those aren’t surprises, but they set the stage for the real question: given who mixes with whom, what does that mean for the early shape of an epidemic?
This is where the contact diaries get translated into transmission through a framework epidemiologists call the next generation matrix. Imagine splitting the population into fifteen slices, five years each, with everyone seventy and older in the last group. For any pair of age classes—say, ten to fourteen-year-olds and thirty-five to thirty-nine-year-olds—you ask: how many at-risk contacts flow from the second group to the first on a typical day?
Those numbers become the elements of a matrix K, with each entry indicating contacts from age group j to age group i. The modeling choice here is to take those entries as proportional to the observed contacts, pooling physical and nonphysical, because pathogen transmissibility is not being estimated—this is about relative incidence by age, not growth rate.
Once you have K, the arithmetic of spread is simple to state, if not simple to compute. Start with a vector x zero that encodes who’s infected in generation zero—one person in a given age class, zeros elsewhere. Multiply by K to get expected new cases in the first generation.
Multiply again for the second generation, and so on: the i-th generation vector, x sub i, equals K raised to the i-th power times x zero. After a few rounds, regardless of where you start, the pattern across age classes converges to the dominant eigenvector of K, the one associated with the largest eigenvalue. Normalize that vector so the entries add up to one, and you’ve got a prediction for the early age distribution of cases in a totally susceptible population.
In their demonstrations, they even set x zero as a single sixty-five to seventy-year-old to show that, after a handful of generations, the age profile of new infections is driven by the contact structure, not the original seed.
So what emerges when you run that machinery on real, smoothed contact matrices? School-age and adolescent groups light up. Across countries, the highest predicted incidence in the initial phase sits in the five to nineteen-year range.
A gentler second peak shows up among adults roughly thirty to forty-four years old, its exact position wobbling by country. That two-humped shape makes sense given what we just walked through. Kids and teens pile up same-age contacts at school and in social circles, pumping fuel into their own age band.
Meanwhile, those secondary diagonals—child-parent contacts around the thirty to forty-year mark—hand the infection back and forth across generations at home.
What’s striking is not just the headline result; it’s the cross-country echo. Italy had many more reported contacts per person than Germany, on average, but once you account for who meets whom, where, and how often, the age-assortative skeleton looks the same. That’s a quiet but powerful message for modelers: the specific numbers will shift, but the geometry of mixing—thick diagonal, fainter family lines, a work-driven adult plateau—may be stable enough to generalize within similar European settings.
Let’s pause for the texture behind those abstractions. When Mossong’s team separated physical from nonphysical contacts, the strong diagonal was still present, but the hot spots got even hotter for school-age and young adult groups. And those long-duration, daily ties that dominate home and close leisure?
They overlap with the places we often go to first with interventions—schools, households, community gatherings. It’s not that workplaces and transit don’t matter; they generate a ton of brief encounters. It’s that the clusters where transmission can really sustain itself often sit where time and touch concentrate.
Of course, diaries aren’t omniscient. They are self-reports from a single day. People forget.
Some countries’ instructions asked respondents to cap frequent professional encounters, which can lop off the long tail of very high-contact jobs. Younger participants were deliberately oversampled. Though every country used a common design, details of recruitment and follow-up differed.
This study never claims to be the voice of Europe as a whole; it claims that a few big patterns recur. There’s also a more technical caveat: the modeling didn’t incorporate network clustering beyond age structure. Real social networks have triangles—your friends are friends with each other—and that can blunt spread compared to a model that only sees average mixing.
All of those caveats pull in the same direction: be careful not to read more precision into the outputs than the inputs warrant.
If you’re curious about the smoothing step I breezed by, here’s what they did. Because raw counts of contacts between, say, twenty-two-year-olds and twenty-seven-year-olds are noisy, they fit a smooth surface over age-by-age cells, using a tensor-product spline with a negative binomial likelihood to respect the overdispersion. That creates a contact matrix that’s faithful to the raw structure but doesn’t overreact to random bumps.
It’s like sharpening a blurry photo without inventing features that aren’t there.
So what does all this buy us? First, an empirical backbone for age-structured transmission models. Instead of plugging in assumed patterns, you can plug in measured ones, with clear uncertainty.
Second, a way to tie setting and intensity to risk in a way that’s portable. We can say, with data, that most high-intensity contacts happen at home and in leisure contexts, that schools are hubs of same-age mixing, and that adults’ work contacts are frequent but often less intimate. That’s the kind of granularity that matters when you’re thinking about which levers move an epidemic early.
I want to underline the discipline of the conclusion. Mossong and colleagues did not declare the best policy or simulate the effect of closing schools versus staggering shifts at work. They showed that if a pathogen spreads through the kinds of contacts they measured, then school-age groups will likely carry the early incidence, with adults in their thirties and early forties as the secondary bridge.
They showed that weekdays are busier than Sundays. They showed that the arithmetic of spread cares deeply about the diagonal in the contact matrix. Everything else—vaccination targeting, testing cadence, social distancing designs—should take those measured facts seriously, but it’s a separate step.
If you’re thinking ahead, there are obvious places to build. The same next-generation framework can accommodate setting-specific transmissibility—home versus work versus school—if you have the data. It can weight physical contact more heavily than nonphysical, again if you can estimate that weight.
And it can be married to mobility and clustering data to soften the assumption that each contact is independent. But those are tomorrow’s models. The contribution here is today’s map: seven thousand two hundred ninety people, ninety-seven thousand nine hundred four encounters, a strong diagonal of age-assortative mixing, and a clear mechanistic bridge from everyday life to the first waves of an epidemic.
Related lectures
- Estimating the infection and case fatality ratio for coronavirus disease (COVID-19) using age-adjusted data from the outbreak on the Diamond Princess cruise ship, February 2020
- Real-time tentative assessment of the epidemiological characteristics of novel coronavirus infections in Wuhan, China, as at 22 January 2020
- Individual Differences in Inhibitory Control, Not Non-Verbal Number Acuity, Correlate with Mathematics Achievement
- The impact of non-pharmaceutical interventions on SARS-CoV-2 transmission across 130 countries and territories
- Community Transmission of Severe Acute Respiratory Syndrome Coronavirus 2, Shenzhen, China, 2020
- High-Resolution Measurements of Face-to-Face Contact Patterns in a Primary School