The Power of Feedback RevisitedA Meta-Analysis of Educational Feedback Research
For more than a decade, the field of educational research carried a confident headline: feedback works, and it works powerfully. Hattie and Timperley's 2007 synthesis reported effect sizes between 0.70 and 0.79. These numbers placed feedback near the very top of every list of effective teaching practices. That claim collected over 25,000 citations. It shaped teacher training, school policy, and classroom culture on several continents. Then Wisniewski, Zierer, and Hattie ran the same test with better tools. The answer came back: the magic is real, but only for certain kinds of feedback. The rest? Close to zero. What you call "feedback" determines almost everything about whether it helps students learn at all. The original Hattie and Timperley synthesis was a meta-synthesis, a combination of other meta-analyses rather than a fresh integration of primary studies. That's not a scandal; it's a legitimate approach. But it meant the conclusions rested on aggregated averages that could drift upward and mask variability. Kluger and DeNisi had already flagged this in 1996, combining one hundred thirty-one studies and over twelve thousand participants to report an average effect of just 0.38, and noting that roughly one third of effects in their data were negative. One third. That number doesn't fit the confident headline, and it had been sitting in the literature for years.
So Wisniewski and colleagues went back to the primary studies. Their final dataset included four hundred thirty-five studies, nine hundred ninety-four effect sizes, and more than sixty-one thousand participants. They used a random-effects model, a statistical framework that doesn't assume every study is measuring the same underlying phenomenon. That matters enormously for something as varied as feedback, which spans simple gold stars in a kindergarten classroom to high-information coaching in elite sports training. A random-effects model treats observed effects as a sample from a distribution of real effects, allowing genuine differences between settings, ages, and feedback formats to show up in the data rather than get averaged away. They also tested six moderator variables: research design, publication type, outcome measure, type of feedback, feedback channel, and feedback direction to explain where that variation came from. The headline number is a Cohen's d of 0.48. A Cohen's d expresses the standardized difference between groups. At 0.48, feedback falls in the medium-effect range, roughly half a standard deviation separating students who receive it from those who don't. That's meaningful. However, the number immediately needs a warning label.
The heterogeneity in this dataset is enormous. The I-squared statistic, which measures what proportion of variance across studies reflects real differences rather than sampling noise, came in at eighty-three percent after removing extreme outliers. That means the average is hiding a distribution that runs from strongly positive effects down to negative ones. Seventeen percent of individual effects were negative. The Q-statistic before trimming was over seven thousand on nine hundred ninety-three degrees of freedom. These aren't rounding errors; they're the signature of a literature where the label "feedback" is being applied to genuinely different things with genuinely different impacts. The overall number is defensible, but it's not actionable. To know what to do, you need to look inside it. Here's what's inside. The single strongest moderator is information content — what the feedback actually tells the learner. The meta-analysis divides feedback into three broad types. Reinforcement or punishment, the simplest and lowest-information class, produces a small effect with a Cohen's d of 0.24. Corrective feedback, telling students whether they got something right or wrong and what the correct answer is, produces a medium effect with a Cohen's d of 0.46. High-information feedback, which supplies task information alongside process guidance and often self-regulation cues, yields a very large effect with a Cohen's d of 0.99, nearly double the overall average.
The more information feedback contains, the more powerful it is. That finding is clean and consistent. Hattie and Timperley's original framework helps explain why. They distinguished three functions of feedback: feed-up, which clarifies the goal; feed-back, which compares current performance to expectations; and feed-forward, which tells the learner what to do next. They also identified four levels at which feedback can operate: task, process, self-regulation, and self. Task-level feedback addresses surface correctness. Process-level feedback addresses the strategies a learner is using. Self-regulation feedback addresses how learners monitor and manage their own approach. Self-level feedback is praise or global judgment about the person. The new meta-analysis confirms this hierarchy empirically: feedback that reaches process and self-regulation levels outperforms task-level feedback, and self-level praise is the least effective of all. The field had theorized this. Now there's a primary-study dataset behind it. The outcome domain matters too. Feedback effects are largest for cognitive outcomes, such as student achievement and retention, at a Cohen's d of 0.51. Motor and physical skills show an even larger effect at a Cohen's d of 0.63.
Motivational outcomes are more modest with a Cohen's d of 0.33, and behavioral outcomes are estimated imprecisely because so few studies measured them. Crucially, many of the negative effects on motivation are tied specifically to low-information feedback — rewards and punishments that can undermine a learner's sense of autonomy and competence. The type of feedback and the type of outcome you're targeting are not independent questions. Now, where does this leave Hattie and Timperley's original framework? The news is mixed in a specific and instructive way. The conceptual map they built — the feed-up, feed-back, feed-forward structure; the four levels of information; the warning about self-level praise — holds up. The new evidence supports all of it. What doesn't hold up is the magnitude of the headline effect. The earlier meta-syntheses reported effects between 0.70 and 0.79. This primary-study meta-analysis lands at 0.48. That gap has a traceable source. The funnel-plot asymmetry test, a diagnostic for publication bias, was highly significant, with an Egger's z of 9.52. When the authors broke this down by publication type, journal articles drove the asymmetry with a z of 9.75, while dissertations showed none at all. Studies with dramatic results are more likely to get published, and that upward pressure on the published literature appears to have inflated the earlier synthetic estimates.
The new work also found that some moderators Hattie and Timperley discussed — particularly timing and valence, the question of whether immediate or delayed feedback is better, and whether positive or negative feedback is better — couldn't be resolved decisively. The primary studies simply didn't provide sufficient or compatible data. Feedback channel, which includes oral, written, and computer-assisted feedback, turned out to be the one moderator that wasn't statistically significant in this analysis. That's a reminder that not every theoretically interesting variable shows up as empirically distinguishable once you look across hundreds of real studies with real noise. What all of this adds up to is a correction that sharpens rather than dismantles. The field was right to focus on feedback. But it was wrong to treat it as one thing. When Wisniewski and colleagues write that "different forms of feedback should be interpreted as independent measures," they're not being diplomatic — they're pointing to a structural error in how the field has talked about the topic. An average effect of 0.48 computed across reinforcement, praise, corrective feedback, and high-information coaching is a bit like computing the average effectiveness of "medicine" across aspirin, chemotherapy, and a placebo. The category is too broad to be useful.
The practical map that emerges is more honest and more useful than the old headline. For cognitive learning goals, process-level and self-regulation-level feedback with clear information value consistently outperforms praise and simple correctness checks. For motivation, the evidence is weaker, and the risks of uninformative reward-and-punishment cycles are real. For researchers, disaggregating feedback type, outcome domain, and information content isn't optional precision; it's the difference between a finding that guides practice and one that merely sounds reassuring. Feedback is worth the focus the field has given it. The evidence is clear on that. What the field now knows, with sixty-one thousand participants behind the claim, is that the question was never whether feedback works. It was always which feedback, for what, and how much it actually tells the learner about where they are and what to do next. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Unternehmen: Warum gründen Frauen seltener?
- Inferring Behavioral Regimes in Urban Mobility via Spatio-Temporal Optimal Transport
- Is volunteering a public health intervention? A systematic review and meta-analysis of the health and survival of volunteers
- Addressing disparities in academic medicine: what of the minority tax?
- AI: A cure for Baumol's disease?
- Eight Americas: Investigating Mortality Disparities across Races, Counties, and Race-Counties in the United States