Feedback sources in essay writingpeer-generated or AI-generated feedback?

Seyyed Kazem Banihashem, Nafiseh Taghizadeh Kerman, Omid Noroozi, Jewoong Moon, Hendrik DrachslerView original
OverviewBalancedjames voice
ChatGPT writes feedback that sounds more complete, more detailed, and more articulate than what peers produce — and yet peer feedback is more useful for actually improving an essay. Hold that paradox for a second. It feels like it shouldn't be true. Better prose, longer comments, more thorough coverage — that should mean better feedback. The study by Banihashem and colleagues ran the experiment to find out exactly why it isn’t, and the answer comes down to a distinction most of us have never consciously made: the difference between describing what is there and identifying what is wrong. The problem that motivates the study is practical and growing. Large university courses have a feedback problem. Teachers simply cannot give individualized, formative feedback to every student at scale — not in a class of fifty, and certainly not in a class of two hundred. Peer feedback became the standard workaround. Students review each other's work, distribute the load, gain practice evaluating arguments, and in theory, everyone benefits. But Banihashem and colleagues are direct about the limitation: for complex tasks like writing an argumentative essay, peers often fall short. The task requires sustained critical thinking, domain knowledge, and the ability to evaluate the coherence of a position and its supporting arguments. Those are exactly the cognitive moves that students are still developing. Without support, peers struggle to generate consistent, diagnostic comments. Enter ChatGPT. After the emergence of large language models capable of reading and responding to student writing, a global argument broke out over whether AI could serve as a new feedback source. The question wasn't just theoretical — it had real consequences for how instructors design courses and how students learn. Banihashem and colleagues set out to answer it empirically. Their study enrolled seventy-four graduate students from a Dutch life sciences program — seventy-seven percent female and twenty-three percent male — in an online Argumentative Essay Writing module. The design ran across two consecutive weeks. In week one, students each wrote an argumentative essay on a course-related controversy: topics included whether scientists with industry affiliations should abstain from risk assessment, and whether food safety is ultimately the consumer's responsibility. In week two, students were randomly assigned to review two peers' essays through a feedback app called FeedbackFruits, with each feedback response constrained to between two hundred fifty and three hundred fifty words. The same essays were also submitted to ChatGPT using the same feedback prompt, with a minor modification for the AI. Essay quality was coded using a scheme developed by Noroozi and colleagues in 2016, which evaluates eight argumentative elements — things like clear stance, pro arguments, counterarguments, and responses to counterarguments — each scored from zero to three. Feedback quality was coded using a 2022 scheme by the same group, which separates feedback into affective, cognitive, and constructive components, with the cognitive dimension further broken into description, identification, and justification. Two trained coders scored both essays and feedback, reaching seventy-five percent agreement with a Cohen's Kappa of zero point seventy-five in both cases. To compare feedback sources, the team used a multivariate analysis of covariance — controlling for gender — and Spearman's correlation to explore links between essay quality and feedback quality. Now for what they found. ChatGPT was significantly better on exactly one dimension: cognitive description. Its mean score was two point zero with a standard deviation of zero — it maxed out every time, with no variation. Peers scored one point ninety-one on the same dimension. That difference was statistically significant, with an effect size of about three percent. What does descriptive feedback actually look like? It summarizes the essay — reports what the writer argued, notes which claims were supported by citations, and offers high-level suggestions framed as general improvements. It reads well. It sounds thorough. ChatGPT produced extended, readable accounts of each essay's form and content. But here is where the paradox resolves. On cognitive identification — the dimension that measures whether feedback actually pinpoints specific problems — peers scored one point fifty-two and ChatGPT scored one point twenty-nine. That gap was also significant, with a p-value below zero point zero one. A typical peer comment read something like: "Since I think your position is missing in the introduction section…" — directly naming a structural omission, telling the writer exactly where to look. ChatGPT's suggestions tended toward elaboration: add real-life examples, strengthen the argument — advice that describes what better writing might look like without diagnosing what is currently broken. There is a real difference between telling someone their house has good bones and pointing out the foundation crack. ChatGPT did more of the former; peers did more of the latter. The other dimensions — affective feedback, cognitive justification, constructive feedback — showed no significant difference between the two sources. Both performed comparably on praise and encouragement, on providing rationales for their comments, and on offering concrete suggestions. The meaningful distinction lived entirely in description versus identification. The correlation findings add another layer. Banihashem and colleagues used Spearman's correlations to test whether better essays attracted better feedback. The intuitive expectation is yes — a stronger essay gives a reviewer more to work with, more nuance to engage. The data said otherwise. Overall essay quality did not predict overall feedback quality from either source. The Spearman rho between overall essay quality and overall ChatGPT feedback was negative zero point zero eight, not significant. For peer feedback, it was negative zero point zero five, also not significant. The more interesting pattern appeared in the affective component. ChatGPT's affective feedback correlated positively with essay quality — the better the essay, the more affective, encouraging language ChatGPT used. The correlation with overall essay quality was zero point twenty-eight. Peers moved in the opposite direction: their affective feedback correlated negatively with essay quality at negative zero point twenty-three. So peers were more encouraging toward weaker essays and more restrained toward stronger ones. ChatGPT was more effusive as essays improved. That is a meaningful behavioral difference, and it has a plausible interpretation — peers may be compensating for weaker writers, softening the criticism, while ChatGPT's tone tracks quality more directly. What does this mean for the classroom? Banihashem and colleagues are clear: neither source wins outright. The data point toward a complementary arrangement. ChatGPT can supply scalable, descriptive coverage consistently — it does not get tired, does not vary, and does not show up to the review with different levels of engagement depending on the day. In large-enrollment courses where teacher time is the binding constraint, that consistency has real value. But peers bring something AI currently underperforms at: they find the actual problem. They read as fellow writers who know the assignment, the stakes, and the argument being made — and they notice when something is missing or broken in a way that ChatGPT, summarizing and describing, does not. The limitations matter and the authors state them honestly. The study drew on seventy-four students from a single institution and a single course. Interrater reliability sat at seventy-five percent — adequate, but improvable. ChatGPT was the only AI tool tested, and the researchers used the same prompt structure for peers and the model despite knowing that AI outputs are sensitive to how prompts are framed. Critically, the study did not track whether students used the feedback during revision — so the chain from feedback quality to learning outcome remains unmeasured. Banihashem and colleagues call for follow-up work at larger scale, with more diverse populations, with prompt validation, with blind assessment designs, and with explicit tracking of how students act on feedback before and after revision. They also point to sequential analysis and data mining as tools that could unpack the dynamics of adaptive feedback in ways that static correlation cannot. The core finding is durable even within those constraints: describing an essay and diagnosing its problems are not the same cognitive act. ChatGPT excels at the former and underperforms at the latter. Peers do the reverse. If you are designing a writing course, that asymmetry is where the design opportunity lives — not in choosing one source over the other, but in building a workflow that uses each for what it actually does well. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

ChatGPT writes feedback that sounds more complete, more detailed, and more articulate than what peers produce — and yet peer feedback is more useful for actually improving an essay. Hold that paradox for a second. It feels like it shouldn't be true. Better prose, longer comments, more thorough coverage — that should mean better feedback. The study by Banihashem and colleagues ran the experiment to find out exactly why it isn’t, and the answer comes down to a distinction most of us have never consciously made: the difference between describing what is there and identifying what is wrong. The problem that motivates the study is practical and growing. Large university courses have a feedback problem. Teachers simply cannot give individualized, formative feedback to every student at scale — not in a class of fifty, and certainly not in a class of two hundred. Peer feedback became the standard workaround. Students review each other's work, distribute the load, gain practice evaluating arguments, and in theory, everyone benefits. But Banihashem and colleagues are direct about the limitation: for complex tasks like writing an argumentative essay, peers often fall short. The task requires sustained critical thinking, domain knowledge, and the ability to evaluate the coherence of a position and its supporting arguments. Those are exactly the cognitive moves that students are still developing. Without support, peers struggle to generate consistent, diagnostic comments.

Enter ChatGPT. After the emergence of large language models capable of reading and responding to student writing, a global argument broke out over whether AI could serve as a new feedback source. The question wasn't just theoretical — it had real consequences for how instructors design courses and how students learn. Banihashem and colleagues set out to answer it empirically. Their study enrolled seventy-four graduate students from a Dutch life sciences program — seventy-seven percent female and twenty-three percent male — in an online Argumentative Essay Writing module. The design ran across two consecutive weeks. In week one, students each wrote an argumentative essay on a course-related controversy: topics included whether scientists with industry affiliations should abstain from risk assessment, and whether food safety is ultimately the consumer's responsibility. In week two, students were randomly assigned to review two peers' essays through a feedback app called FeedbackFruits, with each feedback response constrained to between two hundred fifty and three hundred fifty words. The same essays were also submitted to ChatGPT using the same feedback prompt, with a minor modification for the AI.

Essay quality was coded using a scheme developed by Noroozi and colleagues in 2016, which evaluates eight argumentative elements — things like clear stance, pro arguments, counterarguments, and responses to counterarguments — each scored from zero to three. Feedback quality was coded using a 2022 scheme by the same group, which separates feedback into affective, cognitive, and constructive components, with the cognitive dimension further broken into description, identification, and justification. Two trained coders scored both essays and feedback, reaching seventy-five percent agreement with a Cohen's Kappa of zero point seventy-five in both cases. To compare feedback sources, the team used a multivariate analysis of covariance — controlling for gender — and Spearman's correlation to explore links between essay quality and feedback quality. Now for what they found. ChatGPT was significantly better on exactly one dimension: cognitive description. Its mean score was two point zero with a standard deviation of zero — it maxed out every time, with no variation. Peers scored one point ninety-one on the same dimension. That difference was statistically significant, with an effect size of about three percent. What does descriptive feedback actually look like?

It summarizes the essay — reports what the writer argued, notes which claims were supported by citations, and offers high-level suggestions framed as general improvements. It reads well. It sounds thorough. ChatGPT produced extended, readable accounts of each essay's form and content. But here is where the paradox resolves. On cognitive identification — the dimension that measures whether feedback actually pinpoints specific problems — peers scored one point fifty-two and ChatGPT scored one point twenty-nine. That gap was also significant, with a p-value below zero point zero one. A typical peer comment read something like: "Since I think your position is missing in the introduction section…" — directly naming a structural omission, telling the writer exactly where to look. ChatGPT's suggestions tended toward elaboration: add real-life examples, strengthen the argument — advice that describes what better writing might look like without diagnosing what is currently broken. There is a real difference between telling someone their house has good bones and pointing out the foundation crack. ChatGPT did more of the former; peers did more of the latter.

The other dimensions — affective feedback, cognitive justification, constructive feedback — showed no significant difference between the two sources. Both performed comparably on praise and encouragement, on providing rationales for their comments, and on offering concrete suggestions. The meaningful distinction lived entirely in description versus identification. The correlation findings add another layer. Banihashem and colleagues used Spearman's correlations to test whether better essays attracted better feedback. The intuitive expectation is yes — a stronger essay gives a reviewer more to work with, more nuance to engage. The data said otherwise. Overall essay quality did not predict overall feedback quality from either source. The Spearman rho between overall essay quality and overall ChatGPT feedback was negative zero point zero eight, not significant. For peer feedback, it was negative zero point zero five, also not significant. The more interesting pattern appeared in the affective component. ChatGPT's affective feedback correlated positively with essay quality — the better the essay, the more affective, encouraging language ChatGPT used. The correlation with overall essay quality was zero point twenty-eight.

Peers moved in the opposite direction: their affective feedback correlated negatively with essay quality at negative zero point twenty-three. So peers were more encouraging toward weaker essays and more restrained toward stronger ones. ChatGPT was more effusive as essays improved. That is a meaningful behavioral difference, and it has a plausible interpretation — peers may be compensating for weaker writers, softening the criticism, while ChatGPT's tone tracks quality more directly. What does this mean for the classroom? Banihashem and colleagues are clear: neither source wins outright. The data point toward a complementary arrangement. ChatGPT can supply scalable, descriptive coverage consistently — it does not get tired, does not vary, and does not show up to the review with different levels of engagement depending on the day. In large-enrollment courses where teacher time is the binding constraint, that consistency has real value. But peers bring something AI currently underperforms at: they find the actual problem. They read as fellow writers who know the assignment, the stakes, and the argument being made — and they notice when something is missing or broken in a way that ChatGPT, summarizing and describing, does not. The limitations matter and the authors state them honestly. The study drew on seventy-four students from a single institution and a single course. Interrater reliability sat at seventy-five percent — adequate, but improvable.

ChatGPT was the only AI tool tested, and the researchers used the same prompt structure for peers and the model despite knowing that AI outputs are sensitive to how prompts are framed. Critically, the study did not track whether students used the feedback during revision — so the chain from feedback quality to learning outcome remains unmeasured. Banihashem and colleagues call for follow-up work at larger scale, with more diverse populations, with prompt validation, with blind assessment designs, and with explicit tracking of how students act on feedback before and after revision. They also point to sequential analysis and data mining as tools that could unpack the dynamics of adaptive feedback in ways that static correlation cannot. The core finding is durable even within those constraints: describing an essay and diagnosing its problems are not the same cognitive act. ChatGPT excels at the former and underperforms at the latter. Peers do the reverse. If you are designing a writing course, that asymmetry is where the design opportunity lives — not in choosing one source over the other, but in building a workflow that uses each for what it actually does well. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Social Sciences