Reading the Mind in the Eyes or Reading between the Lines? Theory of Mind Predicts Collective Intelligence Equally Well Online and Face-To-Face
Groups who could see each other's faces and groups who could only read each other's words performed almost identically — not just on one task, but across an entire battery of them. Let that sit for a second. Then consider this: what predicted which groups did well in both conditions wasn't IQ, wasn't personality, and wasn't experience. It was something much stranger — a test that asks you to look at photographs of people's eyes and guess what they are thinking. Woolley and colleagues established something striking about groups: they behave like individuals when it comes to intelligence. Just as a battery of cognitive tasks reveals a single general factor — what we call g — in individuals, the same statistical logic applied to teams reveals a general factor for groups. Woolley and colleagues called it collective intelligence, or c. In their original work, that first factor accounted for forty-three percent of the variance in group performance across a wide range of tasks. That's roughly comparable to what g explains in individual cognitive batteries. And crucially, c predicted how well groups would solve new, more complex tasks later — above and beyond what you would get by averaging the IQ scores of the group's members. That last point matters. Collective intelligence is not a repackaging of individual smarts. Having a few high-IQ members in your group does not guarantee a high-c group.
What did predict c was more interesting: groups with a more equal distribution of conversational turns were more collectively intelligent, and groups with a higher proportion of women tended to score higher on c — an effect largely mediated, Woolley and colleagues found, by members' social perceptiveness. The question that remained open was whether any of this would survive in online environments, where groups communicate through constrained channels and much of what we think of as social information simply isn't available. That's the question Engel, Woolley, Jing, Chabris, and Malone set out to answer. Their key tool for measuring social perceptiveness was the Reading the Mind in the Eyes test — the RME. In this test, participants look at photographs showing only the region around a person's eyes and must infer the emotion or mental state being expressed. It's among the most widely validated adult measures of Theory of Mind, the cognitive capacity to reason about what other people believe, intend, or feel. The RME reliably differentiates neurotypical adults from those on the autism spectrum, shows sensitivity to sex differences and educational background, and in experimental work responds to oxytocin administration and prenatal testosterone exposure. It's a serious, well-established measure.
But here's the puzzle it creates. The test is built around looking at eyes. If what it actually measures is skill at reading facial expressions, then it should lose its predictive power the moment you take away the faces. And in online groups communicating only through text, faces are gone entirely. So either the RME measures something narrower than we thought, and it should stop working — or it measures something deeper, and it won't. To find out, Engel and colleagues recruited two hundred seventy-two adults in the Boston area and organized them into sixty-eight four-person teams. The sample was fifty-two percent male, with an average RME score of twenty-four point eighty-five. Groups were randomly assigned to one of two conditions: face-to-face, where they sat together and could talk freely, or online, where they sat among other participants and could communicate only via text chat.
Both conditions completed the same sixty-four-minute battery of group tasks, delivered through a shared online interface. The tasks were carefully chosen to span different categories of group work: brainstorming problems like generating uses for a brick, intellective tasks including an eighteen-item subset of Raven's Advanced Progressive Matrices and a Sudoku puzzle, judgment tasks requiring predictions about how others would rate slogans, executing tasks like group transcription, remembering tasks drawn from video or complex images, and sensing tasks involving pattern detection in noisy displays. The range was the point — the more varied the tasks, the more meaningful a common factor would be if one emerged. And one did. Factor analysis on the group task scores produced a dominant first factor in both conditions. That factor explained forty-nine percent of the variance for face-to-face groups and forty-one percent for online groups. Collective intelligence is not a face-to-face phenomenon. It exists just as clearly when groups are communicating only by typing words at each other. Now for the finding that shouldn't work. Average group RME scores correlated with collective intelligence at zero point fifty-three in face-to-face groups and zero point fifty-five in online groups. Both significant, both nearly identical.
In regression models predicting collective intelligence from RME alone, the standardized coefficients were zero point fifty-two for face-to-face groups and zero point fifty-five for online groups, with the models explaining twenty-seven percent and thirty-one percent of variance respectively. The group that never saw each other — that operated entirely through text — was predicted just as well by a test built around reading eyes as the group that sat across a table from one another. Personality, by contrast, explained nothing. Engel and colleagues found no significant correlation between a general personality factor and collective intelligence in either condition. Individual cognitive ability alone was, at best, a modest predictor in prior work. The RME was doing something those measures couldn't. Conversational dynamics filled in the rest of the picture. In both conditions, total communication correlated positively with collective intelligence — zero point fifty-two for face-to-face groups and zero point forty-seven for online groups. The distribution of communication went the other direction: groups where one or two people dominated, while others stayed quiet, were less collectively intelligent.
The standard deviation of speaking time or word counts correlated negatively with c at zero point negative thirty for face-to-face and zero point negative forty-one for online groups. Equal participation mattered, and it mattered similarly regardless of whether people were talking or typing. When Engel and colleagues ran regression models that included both RME and the communication variables, the combined models explained forty-one percent of variance for face-to-face groups and forty-two percent for online groups — the two conditions still essentially equivalent. The gender finding also replicated online. The proportion of women in a group predicted collective intelligence in online groups at zero point forty-one, and average RME scores statistically mediated part of that effect, with a Sobel z of one point ninety-five. The mechanism running through social perceptiveness appears to be the same one driving the gender composition effect — online or off. So what is the RME actually measuring? It cannot be merely the ability to read facial expressions, because removing the faces entirely didn't diminish its predictive power at all. Engel and colleagues conclude the RME taps a domain-independent capacity for social reasoning — the ability to infer mental states and intentions from whatever information is available, whether that's the subtle musculature around someone's eyes or the phrasing of a sentence in a chat window.
They put it directly: the RME measures not just the ability to read emotions in eyes but also the ability to "read between the lines" of text-based online interactions. Theory of Mind, in this framing, is portable. It travels across communication channels because it isn't really about the channel at all. That has a practical edge. We have spent considerable energy debating what gets lost when teams go remote — the casual hallway conversation, the ability to read the room, the nonverbal texture of collaboration. Those concerns are real. But this study suggests that the single strongest individual-level predictor of group effectiveness is not something you lose when you log into a chat interface. The capacity to model other minds, to track what your collaborators are thinking and feeling and wanting, operates just as powerfully in a Slack channel as it does in a conference room. Whether that capacity can be trained remains an open question. Engel and colleagues note that some studies suggest temporary improvement on the RME following exposure to literary fiction, but the durability of any such effect is unresolved. What the data do establish is that when you are building a team — or trying to understand why one team outperforms another — the relevant variable is not how smart the individuals are in isolation, but how well they read each other. And that skill, it turns out, doesn't need a face to work with. This lecture was created by ennepō.
Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Collective Cell Motion in an Epithelial Sheet Can Be Quantitatively Described by a Stochastic Interacting Particle Model
- Resolving the gravitational redshift across a millimetre-scale atomic sample
- Existence of chaos for partial difference equations via tangent and cotangent functions
- Constraint on the matter–antimatter symmetry-violating phase in neutrino oscillations
- Temporal Patterns of Happiness and Information in a Global Social Network: Hedonometrics and Twitter
- What Is Stochastic Resonance? Definitions, Misconceptions, Debates, and Its Relevance to Biology