The Ordinal Effects of OstracismA Meta-Analysis of 120 Cyberball Studies

Chris Hartgerink, Ilja van Beest, Jelte M. Wicherts, Kipling D. WilliamsView original
OverviewBalancedadam voice
For decades, social psychologists knew that ostracism hurt. You could see it in people's faces and feel it in a room when someone was being deliberately ignored. But knowing something hurts and measuring exactly how much it hurts — across hundreds of studies, thousands of people, with a number you can actually hold up and examine — those are entirely different things. A meta-analysis by Chris Hartgerink, Ilja van Beest, Jelte Wicherts, and Kipling Williams finally produced that number. And it's larger than almost anyone expected. The tool that made all of this possible is called Cyberball. It is, on the surface, absurdly simple. Participants sit down and are told that they're going to play a virtual ball-tossing game with two or three other people connected online. They see a little animated figure on screen representing themselves and two or more figures representing the other players. The ball gets tossed around. And then, after a couple of passes, it stops coming to you. That's it. The other players just stop including you. What the participant doesn't know is that the other players are controlled by a computer — there are no other real people. The exclusion is programmed. The pain, however, is entirely real. What makes Cyberball so powerful is precisely that minimal deception. It doesn't announce rejection. It doesn't tell you that you're disliked or unwanted. It simply lets inclusion stop, quietly, mid-game. Hartgerink and colleagues note that this mirrors something genuinely common — research suggests that most people are ignored or excluded at least once a day. Cyberball captures that ordinary, low-drama form of being left out and brings it into the lab. The literature it generated is enormous: the authors identified more than 200 published papers using the paradigm, with over 19,500 participants. Their meta-analysis zeroed in on 120 of those studies, carefully selected from 98 papers, with a combined sample of 11,869 participants. To make sense of what all those studies were testing, you need to understand the theory they were built around. Kipling Williams — one of the meta-analysis authors and the originator of Cyberball — developed what he called the temporal need-threat model. The idea is that ostracism unfolds in stages. The first is reflexive: immediate, automatic, visceral. Being excluded threatens four fundamental needs — belonging, self-esteem, control, and meaningful existence. Williams argued that this initial reaction is essentially a social reflex. It doesn't pause to assess context. It doesn't ask whether the people ignoring you are strangers, enemies, or people whose opinion you care about. It just hurts, immediately and reliably. Then comes what Williams called the reflective stage. This is where deliberate cognitive processing kicks in. People start to make sense of what happened, reinterpret it, cope with it. And here, Williams predicted, context should start to matter. The type of group excluding you, your individual personality, and the perceived intentionality of the exclusion — these factors should shape how much the exclusion stings at this later stage. The core testable prediction was clean: immediate effects are hard to budge, delayed effects are more variable. The meta-analysis was built to put that prediction to the test. Before getting to the results, it's worth appreciating how carefully this meta-analysis was constructed. The literature search used seven complementary strategies — database searches, citation tracking, conference abstracts, personal communications, and a direct call for unpublished data. From two thousand four hundred and sixty-eight potentially relevant records, the team eventually arrived at 120 independent Cyberball experiments. Effect sizes were calculated as Hedges's g — the bias-corrected version of the standardized mean difference, which adjusts for the slight overestimation that creeps into Cohen's d with smaller samples. Crucially, they extracted two effect sizes from each study: one for the first dependent measure collected after the manipulation and one for the last. This first versus last distinction was their operationalization of Williams's reflexive versus reflective distinction. They preregistered their analyses and posted the full data package publicly. The statistical models were random-effects and mixed-effects, with sensitivity checks, outlier diagnostics, and funnel plots to assess publication bias. Now, the findings. The average immediate ostracism effect is larger than 1.4 standard deviations. Hartgerink and colleagues estimated the first measure effect at a value of negative 1.36, with a ninety-five percent confidence interval running from negative 1.54 to negative 1.18. That is a large effect by any standard in psychology. It means the average ostracized participant feels substantially worse than the vast majority of included participants on the same measures. And this effect doesn't fragment when you look at different versions of the paradigm. Mixed-effects models found no strong moderation by country, participant age, proportion of female participants, number of ball tosses, or game duration. In a tighter homogeneous subset — studies using three players, thirty throws, and under five minutes — the effect was even larger, at a value of negative 2.05. Being briefly excluded from a trivial game with strangers produces one of the most consistent large effects in social psychology. The effect does attenuate over time within studies. By the last measure — on average about four and a half minutes after the game — the effect drops to approximately negative 0.76, with a confidence interval of negative 0.86 to negative 0.59. Still substantial, but meaningfully smaller than the immediate effect. The confidence intervals for first and last measures don't overlap. The effect also varies somewhat by type of outcome measure. Intrapersonal measures and fundamental needs — belonging, control, meaningful existence, and self-esteem — show robust negative effects. Interpersonal measures, things like behavioral intentions toward others, show somewhat weaker and more variable effects. Here is where the theory runs into trouble. Williams predicted that the immediate, reflexive effects would be essentially resistant to moderation — context wouldn't much change how hard that first moment of exclusion hits. The meta-analysis found something different. Across the fifty-two factorial studies that included a moderating variable, the average interaction effect on the first measure was negative 0.46, with a p-value below 0.001 and a confidence interval of negative 0.64 to negative 0.28. Moderators do shift the size of the immediate ostracism effect. The reflexive stage is not as impervious as Williams proposed. And the story doesn't straighten out at the delayed measure either. The estimated interaction effect on the last measure was negative 0.20, which just missed significance at a p-value of 0.052. More telling, when the authors asked whether time elapsed since the exclusion predicted how much moderation appeared at the last measure, the answer was no. The regression coefficient for time was essentially zero and non-significant, even in sensitivity analyses that removed outliers. More time after the ball stops coming doesn't mean more contextual factors shaping how much you hurt. The reflective stage doesn't appear to simply switch on with the passage of minutes. Hartgerink and colleagues describe this pattern as an ordinal effect. The direction of ostracism's impact is consistent — it always produces worse outcomes for excluded participants than for included ones. But the magnitude of that effect varies by context even at the earliest measurement, and that variation doesn't follow the temporal logic Williams predicted. This is a refinement, not a demolition, of the model. The core architecture — that ostracism threatens fundamental needs, that those effects are immediate and broad — holds up. The sharp division between a moderation-immune reflexive stage and a moderation-susceptible reflective stage does not. What the meta-analysis also makes clear is that something remains unexplained. Even after accounting for all the coded moderators, the residual heterogeneity across studies remains large — an I-squared of roughly ninety-three percent on first-measure effects, meaning the known moderators explain only a small fraction of the variation between studies. There are likely moderators nobody has systematically studied yet. The authors invite other researchers to dig into the publicly available dataset and to run new studies with larger samples. Current Cyberball studies are typically powered to detect the main effect comfortably, but interaction effects are considerably smaller and would require substantially bigger samples to test reliably. Leave the last number where it belongs: roughly 1.5 standard deviations. That's what being left out of a virtual ball-tossing game does, on average, to a person's sense of belonging and self-worth. Not a clinical event, not a humiliation in front of a crowd — just a few minutes of passes that stop coming. The magnitude of that is the finding worth carrying with you. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

For decades, social psychologists knew that ostracism hurt. You could see it in people's faces and feel it in a room when someone was being deliberately ignored. But knowing something hurts and measuring exactly how much it hurts — across hundreds of studies, thousands of people, with a number you can actually hold up and examine — those are entirely different things. A meta-analysis by Chris Hartgerink, Ilja van Beest, Jelte Wicherts, and Kipling Williams finally produced that number. And it's larger than almost anyone expected. The tool that made all of this possible is called Cyberball. It is, on the surface, absurdly simple. Participants sit down and are told that they're going to play a virtual ball-tossing game with two or three other people connected online. They see a little animated figure on screen representing themselves and two or more figures representing the other players. The ball gets tossed around. And then, after a couple of passes, it stops coming to you. That's it. The other players just stop including you. What the participant doesn't know is that the other players are controlled by a computer — there are no other real people. The exclusion is programmed. The pain, however, is entirely real. What makes Cyberball so powerful is precisely that minimal deception. It doesn't announce rejection. It doesn't tell you that you're disliked or unwanted.

It simply lets inclusion stop, quietly, mid-game. Hartgerink and colleagues note that this mirrors something genuinely common — research suggests that most people are ignored or excluded at least once a day. Cyberball captures that ordinary, low-drama form of being left out and brings it into the lab. The literature it generated is enormous: the authors identified more than 200 published papers using the paradigm, with over 19,500 participants. Their meta-analysis zeroed in on 120 of those studies, carefully selected from 98 papers, with a combined sample of 11,869 participants. To make sense of what all those studies were testing, you need to understand the theory they were built around. Kipling Williams — one of the meta-analysis authors and the originator of Cyberball — developed what he called the temporal need-threat model. The idea is that ostracism unfolds in stages. The first is reflexive: immediate, automatic, visceral. Being excluded threatens four fundamental needs — belonging, self-esteem, control, and meaningful existence. Williams argued that this initial reaction is essentially a social reflex. It doesn't pause to assess context. It doesn't ask whether the people ignoring you are strangers, enemies, or people whose opinion you care about. It just hurts, immediately and reliably. Then comes what Williams called the reflective stage. This is where deliberate cognitive processing kicks in. People start to make sense of what happened, reinterpret it, cope with it.

And here, Williams predicted, context should start to matter. The type of group excluding you, your individual personality, and the perceived intentionality of the exclusion — these factors should shape how much the exclusion stings at this later stage. The core testable prediction was clean: immediate effects are hard to budge, delayed effects are more variable. The meta-analysis was built to put that prediction to the test. Before getting to the results, it's worth appreciating how carefully this meta-analysis was constructed. The literature search used seven complementary strategies — database searches, citation tracking, conference abstracts, personal communications, and a direct call for unpublished data. From two thousand four hundred and sixty-eight potentially relevant records, the team eventually arrived at 120 independent Cyberball experiments. Effect sizes were calculated as Hedges's g — the bias-corrected version of the standardized mean difference, which adjusts for the slight overestimation that creeps into Cohen's d with smaller samples. Crucially, they extracted two effect sizes from each study: one for the first dependent measure collected after the manipulation and one for the last. This first versus last distinction was their operationalization of Williams's reflexive versus reflective distinction.

They preregistered their analyses and posted the full data package publicly. The statistical models were random-effects and mixed-effects, with sensitivity checks, outlier diagnostics, and funnel plots to assess publication bias. Now, the findings. The average immediate ostracism effect is larger than 1.4 standard deviations. Hartgerink and colleagues estimated the first measure effect at a value of negative 1.36, with a ninety-five percent confidence interval running from negative 1.54 to negative 1.18. That is a large effect by any standard in psychology. It means the average ostracized participant feels substantially worse than the vast majority of included participants on the same measures. And this effect doesn't fragment when you look at different versions of the paradigm. Mixed-effects models found no strong moderation by country, participant age, proportion of female participants, number of ball tosses, or game duration. In a tighter homogeneous subset — studies using three players, thirty throws, and under five minutes — the effect was even larger, at a value of negative 2.05. Being briefly excluded from a trivial game with strangers produces one of the most consistent large effects in social psychology.

The effect does attenuate over time within studies. By the last measure — on average about four and a half minutes after the game — the effect drops to approximately negative 0.76, with a confidence interval of negative 0.86 to negative 0.59. Still substantial, but meaningfully smaller than the immediate effect. The confidence intervals for first and last measures don't overlap. The effect also varies somewhat by type of outcome measure. Intrapersonal measures and fundamental needs — belonging, control, meaningful existence, and self-esteem — show robust negative effects. Interpersonal measures, things like behavioral intentions toward others, show somewhat weaker and more variable effects. Here is where the theory runs into trouble. Williams predicted that the immediate, reflexive effects would be essentially resistant to moderation — context wouldn't much change how hard that first moment of exclusion hits. The meta-analysis found something different. Across the fifty-two factorial studies that included a moderating variable, the average interaction effect on the first measure was negative 0.46, with a p-value below 0.001 and a confidence interval of negative 0.64 to negative 0.28. Moderators do shift the size of the immediate ostracism effect. The reflexive stage is not as impervious as Williams proposed.

And the story doesn't straighten out at the delayed measure either. The estimated interaction effect on the last measure was negative 0.20, which just missed significance at a p-value of 0.052. More telling, when the authors asked whether time elapsed since the exclusion predicted how much moderation appeared at the last measure, the answer was no. The regression coefficient for time was essentially zero and non-significant, even in sensitivity analyses that removed outliers. More time after the ball stops coming doesn't mean more contextual factors shaping how much you hurt. The reflective stage doesn't appear to simply switch on with the passage of minutes. Hartgerink and colleagues describe this pattern as an ordinal effect. The direction of ostracism's impact is consistent — it always produces worse outcomes for excluded participants than for included ones. But the magnitude of that effect varies by context even at the earliest measurement, and that variation doesn't follow the temporal logic Williams predicted. This is a refinement, not a demolition, of the model. The core architecture — that ostracism threatens fundamental needs, that those effects are immediate and broad — holds up. The sharp division between a moderation-immune reflexive stage and a moderation-susceptible reflective stage does not.

What the meta-analysis also makes clear is that something remains unexplained. Even after accounting for all the coded moderators, the residual heterogeneity across studies remains large — an I-squared of roughly ninety-three percent on first-measure effects, meaning the known moderators explain only a small fraction of the variation between studies. There are likely moderators nobody has systematically studied yet. The authors invite other researchers to dig into the publicly available dataset and to run new studies with larger samples. Current Cyberball studies are typically powered to detect the main effect comfortably, but interaction effects are considerably smaller and would require substantially bigger samples to test reliably. Leave the last number where it belongs: roughly 1.5 standard deviations. That's what being left out of a virtual ball-tossing game does, on average, to a person's sense of belonging and self-worth. Not a clinical event, not a humiliation in front of a crowd — just a few minutes of passes that stop coming. The magnitude of that is the finding worth carrying with you. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Psychology