Influence of COVID-19 confinement on students’ performance in higher education

Teresa González, M.A. de la Rubia, Kajetan Piotr Hincz, M. Comas-Lopez, Laia Subirats, Santi Fort, G. M. SachaView original
OverviewBalancedjames voice
Picture a student's study calendar from before March 2020. There are long stretches of blank space, a few scattered attempts at engagement, and then — right before the deadline — a frantic cluster of activity. That pattern, repeated across hundreds of students, was the baseline. González and colleagues at Universidad Autónoma de Madrid looked at what happened to that calendar when confinement forced students home. They found that one of the most stubborn habits in education quietly broke. The study followed four hundred fifty-eight students across three subjects: Applied Computing, Metabolism, and Design of Water Treatment Facilities. The researchers recorded every self-evaluation attempt through e-valUAM, an online platform that implements computer adaptive testing. This means the difficulty of each question adjusts based on a student's running track record of correct answers. This gave the team something rare: a granular, timestamped log of when students actually chose to study, not just whether they passed. What the control group data showed was blunt. In Applied Computing during twenty seventeen to twenty eighteen, students used the platform six hundred eighty-eight times in May alone, compared to just four hundred eighty-six attempts across the entire preceding three months of February, March, and April. More than thirty-three percent of all self-evaluation tests happened in the final six days before the exam. On the day before the exam: one hundred twenty-seven attempts. On exam day itself: one hundred seventy-three attempts, including seventy-eight that were the exam. The studying wasn't spread across the semester. It was a last-minute sprint. This pattern held across both control years — twenty seventeen to twenty eighteen and twenty eighteen to twenty nineteen — with near-identical results. The Design of Water Treatment Facilities course showed the same thing from a different angle: even when self-evaluation material was made available three weeks before the deadline, students used it almost entirely in the final two days. The habit was persistent. The question was whether anything could actually break it. The experimental group is the twenty nineteen to twenty twenty cohort — the students whose face-to-face instruction was interrupted in March twenty twenty. González and colleagues compared these students against the two prior years' cohorts, and they built their design carefully around a specific concern: maybe any score improvements were just an artifact of easier online tests. To test against that, they chose subjects with different instructor responses. In Applied Computing, instructors increased the number of autonomous assessment tasks during confinement — the count rose from two hundred twenty-five and two hundred forty-six in the two control years to three hundred twenty-eight in twenty nineteen to twenty twenty. In Metabolism, instructors did not add tasks; instead, face-to-face lectures became recorded videos, and many assessment activities, which had always been online, stayed exactly as they were. That second case — Metabolism, unchanged format — is where the argument gets its sharpest edge. Scores in Applied Computing rose after confinement. Before confinement, the twenty nineteen to twenty twenty mean was four point five, compared to three point nine in both prior years — already statistically significant, with a p-value below zero point zero zero three and below zero point zero zero two respectively. After confinement, the gap widened: six point three versus five point six and five point two in the control years, with p-values of zero point zero one six eight and zero point zero zero two. But a skeptic could argue the new tasks were simply easier online, or that more tasks gave students more practice. Metabolism closes that argument. Activity ten — always online, format unchanged — rose from means of six point five and six point seven in the control years to eight point one in twenty nineteen to twenty twenty, both comparisons significant at p-values below zero point zero zero one and zero point zero zero five. Activity eleven went from six point eight and six point one to seven point eight, with p-values of zero point zero zero one four and below zero point zero zero zero one. Pass rates followed: ninety-five point six percent passed activity ten in twenty nineteen to twenty twenty, compared to significantly lower rates in both prior years. For activity eleven, the pass rate hit ninety-seven point seven percent. Same questions, same format, same platform. The scores still went up. The researchers used nonparametric methods throughout — Mann-Whitney and Kruskal-Wallis for group comparisons, Wilcoxon signed-rank for related samples — because the data consistently failed normality checks. So if it wasn't easier tests, what was it? This is where the timestamped data becomes decisive. González and colleagues weren't just measuring outcomes — they were measuring behavior in time. In the Design of Water Treatment Facilities longitudinal data, the team tracked what happened when material was made available earlier and a small reward was added for using the application on three or more different days. The average span of application use increased by twenty-three percent, from one point nine days to two point four days. Attempts per student dropped from nine point five to five point four. Fewer intensive cramming sessions, more distributed logins spread over several days. The Wilcoxon signed-rank test returned a statistic of fifteen point five against a critical value of twenty-five — statistically significant. In Applied Computing, the same behavioral shift appears in the twenty nineteen to twenty twenty data. Normalized tests per student before confinement were two point thirty-one and three point thirty-seven in the control years, rising to three point six in twenty nineteen to twenty twenty. Total tests per student across the full control years were thirteen point eighty-seven and sixteen point one — a difference of two point twenty-three tests per student accumulated over three and a half months. The twenty nineteen to twenty twenty cohort was already engaging more before confinement began. Then engagement and scores both jumped after March twenty twenty. The authors' interpretation is direct: confinement changed the structure of students' days. Without commutes, without the physical rhythm of campus life and social schedules, students shifted toward more frequent, lower-intensity study sessions distributed across the semester rather than concentrated before deadlines. That behavioral shift — from cramming to continuity — is what González and colleagues identify as the mechanism behind the score improvements. The crisis accidentally ran a clean experiment that educational researchers had struggled to design on purpose. There are real limits to acknowledge. This is one institution, three subjects, and a psychological moment that will not repeat exactly. Lockdown brought stress and disruption alongside any study-habit benefit, and the authors make no claim to have disentangled those factors. The unusual conditions of twenty twenty are part of the experimental context, not incidental to it. But the practical signal is clear, and it doesn't depend on a pandemic to act on. The longitudinal Design course showed that even a modest structural intervention — making material available earlier and adding a small reward for multi-day use — produced the same directional shift toward distributed study that confinement produced at scale. Structured self-assessment tools, incentives for spreading out engagement, and adaptive platforms that provide ongoing feedback are the levers the data point to. Students don't naturally study continuously. But when the conditions are right — whether by circumstance or by design — they can. The deepest finding here isn't really about COVID. It's about the gap between how students study when left to their own defaults and how they study when the environment shifts. That gap, measured in test scores across hundreds of students and two years of controls, turns out to be large. And it turns out to be closeable. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Picture a student's study calendar from before March 2020. There are long stretches of blank space, a few scattered attempts at engagement, and then — right before the deadline — a frantic cluster of activity. That pattern, repeated across hundreds of students, was the baseline. González and colleagues at Universidad Autónoma de Madrid looked at what happened to that calendar when confinement forced students home. They found that one of the most stubborn habits in education quietly broke. The study followed four hundred fifty-eight students across three subjects: Applied Computing, Metabolism, and Design of Water Treatment Facilities. The researchers recorded every self-evaluation attempt through e-valUAM, an online platform that implements computer adaptive testing. This means the difficulty of each question adjusts based on a student's running track record of correct answers. This gave the team something rare: a granular, timestamped log of when students actually chose to study, not just whether they passed. What the control group data showed was blunt. In Applied Computing during twenty seventeen to twenty eighteen, students used the platform six hundred eighty-eight times in May alone, compared to just four hundred eighty-six attempts across the entire preceding three months of February, March, and April. More than thirty-three percent of all self-evaluation tests happened in the final six days before the exam.

On the day before the exam: one hundred twenty-seven attempts. On exam day itself: one hundred seventy-three attempts, including seventy-eight that were the exam. The studying wasn't spread across the semester. It was a last-minute sprint. This pattern held across both control years — twenty seventeen to twenty eighteen and twenty eighteen to twenty nineteen — with near-identical results. The Design of Water Treatment Facilities course showed the same thing from a different angle: even when self-evaluation material was made available three weeks before the deadline, students used it almost entirely in the final two days. The habit was persistent. The question was whether anything could actually break it. The experimental group is the twenty nineteen to twenty twenty cohort — the students whose face-to-face instruction was interrupted in March twenty twenty. González and colleagues compared these students against the two prior years' cohorts, and they built their design carefully around a specific concern: maybe any score improvements were just an artifact of easier online tests. To test against that, they chose subjects with different instructor responses.

In Applied Computing, instructors increased the number of autonomous assessment tasks during confinement — the count rose from two hundred twenty-five and two hundred forty-six in the two control years to three hundred twenty-eight in twenty nineteen to twenty twenty. In Metabolism, instructors did not add tasks; instead, face-to-face lectures became recorded videos, and many assessment activities, which had always been online, stayed exactly as they were. That second case — Metabolism, unchanged format — is where the argument gets its sharpest edge. Scores in Applied Computing rose after confinement. Before confinement, the twenty nineteen to twenty twenty mean was four point five, compared to three point nine in both prior years — already statistically significant, with a p-value below zero point zero zero three and below zero point zero zero two respectively. After confinement, the gap widened: six point three versus five point six and five point two in the control years, with p-values of zero point zero one six eight and zero point zero zero two. But a skeptic could argue the new tasks were simply easier online, or that more tasks gave students more practice. Metabolism closes that argument. Activity ten — always online, format unchanged — rose from means of six point five and six point seven in the control years to eight point one in twenty nineteen to twenty twenty, both comparisons significant at p-values below zero point zero zero one and zero point zero zero five.

Activity eleven went from six point eight and six point one to seven point eight, with p-values of zero point zero zero one four and below zero point zero zero zero one. Pass rates followed: ninety-five point six percent passed activity ten in twenty nineteen to twenty twenty, compared to significantly lower rates in both prior years. For activity eleven, the pass rate hit ninety-seven point seven percent. Same questions, same format, same platform. The scores still went up. The researchers used nonparametric methods throughout — Mann-Whitney and Kruskal-Wallis for group comparisons, Wilcoxon signed-rank for related samples — because the data consistently failed normality checks. So if it wasn't easier tests, what was it? This is where the timestamped data becomes decisive. González and colleagues weren't just measuring outcomes — they were measuring behavior in time. In the Design of Water Treatment Facilities longitudinal data, the team tracked what happened when material was made available earlier and a small reward was added for using the application on three or more different days. The average span of application use increased by twenty-three percent, from one point nine days to two point four days. Attempts per student dropped from nine point five to five point four. Fewer intensive cramming sessions, more distributed logins spread over several days.

The Wilcoxon signed-rank test returned a statistic of fifteen point five against a critical value of twenty-five — statistically significant. In Applied Computing, the same behavioral shift appears in the twenty nineteen to twenty twenty data. Normalized tests per student before confinement were two point thirty-one and three point thirty-seven in the control years, rising to three point six in twenty nineteen to twenty twenty. Total tests per student across the full control years were thirteen point eighty-seven and sixteen point one — a difference of two point twenty-three tests per student accumulated over three and a half months. The twenty nineteen to twenty twenty cohort was already engaging more before confinement began. Then engagement and scores both jumped after March twenty twenty. The authors' interpretation is direct: confinement changed the structure of students' days. Without commutes, without the physical rhythm of campus life and social schedules, students shifted toward more frequent, lower-intensity study sessions distributed across the semester rather than concentrated before deadlines. That behavioral shift — from cramming to continuity — is what González and colleagues identify as the mechanism behind the score improvements. The crisis accidentally ran a clean experiment that educational researchers had struggled to design on purpose.

There are real limits to acknowledge. This is one institution, three subjects, and a psychological moment that will not repeat exactly. Lockdown brought stress and disruption alongside any study-habit benefit, and the authors make no claim to have disentangled those factors. The unusual conditions of twenty twenty are part of the experimental context, not incidental to it. But the practical signal is clear, and it doesn't depend on a pandemic to act on. The longitudinal Design course showed that even a modest structural intervention — making material available earlier and adding a small reward for multi-day use — produced the same directional shift toward distributed study that confinement produced at scale. Structured self-assessment tools, incentives for spreading out engagement, and adaptive platforms that provide ongoing feedback are the levers the data point to. Students don't naturally study continuously. But when the conditions are right — whether by circumstance or by design — they can. The deepest finding here isn't really about COVID. It's about the gap between how students study when left to their own defaults and how they study when the environment shifts. That gap, measured in test scores across hundreds of students and two years of controls, turns out to be large. And it turns out to be closeable. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field.

Read when you can. Listen when you want to.

More in Computer Science