The Hawthorne Effecta randomised, controlled trial

Rob McCarney, James Warner, Steve Iliffe, Robbert van Haselen, Mark Griffin, Peter FisherView original
OverviewBalancedriya_rao voice
Everyone in clinical research knows about the Hawthorne Effect. It exists in methods sections the way a ghost lives in a house — everyone acknowledges it, but nobody has actually measured it. The term traces back to worker productivity studies at the Western Electric Company's Hawthorne Works in Chicago in the 1920s, where investigators noticed that nearly any change in working conditions seemed to boost output. The explanation that stuck was psychological: workers improved because they felt singled out, observed, and made to feel important. Over time, the label migrated into medicine, where it came to mean something specific and troubling — that patients in clinical trials may get better simply because they are being studied, not because the treatment works. McCarney and colleagues knew this had been discussed for decades. What they noticed was that no one had ever run a randomised controlled trial to actually quantify it. So they did. The vehicle was a community-based dementia trial called DIGGER, which stands for the Dementia In General Practice Ginkgo Extract Research study. This was a placebo-controlled trial of high-purity Ginkgo biloba at 120 milligrams daily. But embedded within that drug question was a second, independent randomisation. All 176 participants were also assigned to one of two follow-up schedules. The intensive group received comprehensive home assessments at baseline as well as at two, four, and six months after randomisation. The minimal group had an abbreviated baseline visit and a single full assessment at six months. Everything else — medication, placebo, the disease — was the same. The follow-up intensity was the only variable being changed. Participants were recruited mainly through primary care: 119 general practices and 388 individual general practitioners, with three-quarters of participants referred directly by their family doctor. The mean age was seventy-nine point five years. The mean baseline score on the Alzheimer's Disease Assessment Scale cognitive subscale, a seventy-point scale where higher scores mean worse cognition, was twenty-two point seven. The elegance of the design is worth pausing on. The same one hundred seventy-six people answered both questions simultaneously. The drug comparison was double-blind. The follow-up comparison couldn't be blinded — you can't hide from a participant whether researchers are visiting every two months or just once — but it was randomised, which is the most important aspect. After six months, more attention had produced a measurable cognitive benefit. In the primary intention-to-treat analysis, adjusting for baseline score, participants in the intensive follow-up group scored about two points better on the Alzheimer's Disease Assessment Scale cognitive subscale than those in the minimal group. The adjusted mean difference was minus two point zero two, with a ninety-five percent confidence interval ranging from minus three point nine one to minus zero point twelve, and a p-value of zero point zero three seven. The same direction held in the per-protocol analysis, where the difference was slightly larger at minus two point three eight. Two points is modest. McCarney and colleagues say so plainly: "This may not be regarded as particularly clinically significant." But they immediately note it is of similar magnitude to effects reported in randomised controlled trials of cholinesterase inhibitors, the standard drug class for Alzheimer's disease. The observation alone, in other words, produced a cognitive shift comparable to what some treatments claim. Then came the twist. Participant-rated quality of life — measured with the Quality of Life in Alzheimer's Disease, a thirteen-item scale where higher scores mean better quality of life — told a different story entirely. It favoured the minimal follow-up group. The adjusted mean difference was minus one point three eight, with a ninety-five percent confidence interval of minus two point six four to minus zero point twelve, and a p-value of zero point zero three two. Patients who were seen less frequently felt slightly better about their lives. Carer-rated quality of life showed no significant difference either way. So the trial produced a genuine paradox. More intensive follow-up improved measured cognition but slightly worsened how patients rated their own well-being. The cognitive benefit and the quality-of-life cost pointed in opposite directions, and the data alone can't fully resolve why. McCarney and colleagues work through the candidates honestly. The most obvious explanation for the cognitive finding is a practice or learning effect — patients who take the Alzheimer's Disease Assessment Scale cognitive subscale four times may simply get better at the test. The researchers tried to limit this by rotating the word lists used in the memory items. They also note the test-retest coefficient for the Alzheimer's Disease Assessment Scale cognitive subscale over six weeks is zero point nine one, which they cite as evidence against a large learning effect. But they admit they cannot rule it out, particularly in a dementia population where test familiarity and genuine cognitive change are hard to separate. A second candidate is social contact itself: more frequent visits from researchers may improve mood, motivation, or general engagement in ways that show up on cognitive tasks. A third is what the paper calls better recognition of needs — more contact may help participants or families identify and address problems they would otherwise miss. The quality-of-life finding points toward a different mechanism. More frequent reminders of being in a dementia study may increase awareness of the diagnosis and disability, lowering self-rated well-being even as it improves cognitive performance. The burden of repeated assessments — each taking between ninety minutes and two hours — may also weigh on participants in the intensive group in ways that do not register on a cognitive scale but do show up in how someone answers questions about their own life. Notably, carer-rated quality of life was unaffected, which suggests that the effect is happening in how patients perceive and report their own experience, not in anything their carers observe from the outside. The authors are candid about one more complication: adherence to the protocol differed between groups. Nearly thirty percent of participants in the intensive group were not followed exactly according to protocol, compared to about fourteen percent in the minimal group. That likely means the true effect of follow-up intensity is underestimated, not inflated. What the DIGGER trial ultimately demonstrates has reach well beyond this particular study of dementia and Ginkgo biloba. If two points of cognitive improvement can emerge from observation alone — from the act of being followed closely, assessed repeatedly, and contacted by researchers — then the treatment effects reported in rigorous clinical trials may systematically overestimate what happens in routine clinical care. Most trials cannot isolate this component, because intensive monitoring typically applies equally to treatment and placebo arms. It gets subtracted out of the drug comparison but inflates both groups relative to what ordinary patients experience when they see their doctor twice a year. As McCarney and colleagues put it, such an effect "may result in an inflated estimate of effect size in routine clinical settings by over-estimating response in both groups." They raise a provocative corollary. If participation in research itself produces benefit, one response would be to deliver routine care under conditions more similar to trials — closer monitoring and more frequent contact. The authors acknowledge this would have resource implications, then leave it there. But the logic holds: if the observation is therapeutic, the observation has value. The paper also notes that people with dementia may actually show a smaller Hawthorne Effect than cognitively intact populations, because their capacity for learning and social response may be reduced. The two-point Alzheimer's Disease Assessment Scale cognitive subscale shift observed here could be a floor, not a ceiling, for what this phenomenon does in other trials. Before this study, the Hawthorne Effect in clinical research was exactly what the original Hawthorne research had always been — a compelling story with contested evidence behind it. McCarney and colleagues ran the first randomised controlled trial designed to quantify it and found a small but real signal. Two points on a cognitive scale in community-dwelling patients with mild-to-moderate dementia. Modest, as the authors say. But it existed, it was measurable, and it showed up in both the cognitive data and the quality-of-life data, pointing in different directions at once. The effect that spent decades as a cautionary footnote now has a number. Whether it comes from observation, practice, social contact, or something we haven't named yet — the paper is clear that it "remains intriguing and clinically relevant, whatever the cause." This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Everyone in clinical research knows about the Hawthorne Effect. It exists in methods sections the way a ghost lives in a house — everyone acknowledges it, but nobody has actually measured it. The term traces back to worker productivity studies at the Western Electric Company's Hawthorne Works in Chicago in the 1920s, where investigators noticed that nearly any change in working conditions seemed to boost output. The explanation that stuck was psychological: workers improved because they felt singled out, observed, and made to feel important. Over time, the label migrated into medicine, where it came to mean something specific and troubling — that patients in clinical trials may get better simply because they are being studied, not because the treatment works. McCarney and colleagues knew this had been discussed for decades. What they noticed was that no one had ever run a randomised controlled trial to actually quantify it. So they did. The vehicle was a community-based dementia trial called DIGGER, which stands for the Dementia In General Practice Ginkgo Extract Research study. This was a placebo-controlled trial of high-purity Ginkgo biloba at 120 milligrams daily. But embedded within that drug question was a second, independent randomisation.

All 176 participants were also assigned to one of two follow-up schedules. The intensive group received comprehensive home assessments at baseline as well as at two, four, and six months after randomisation. The minimal group had an abbreviated baseline visit and a single full assessment at six months. Everything else — medication, placebo, the disease — was the same. The follow-up intensity was the only variable being changed. Participants were recruited mainly through primary care: 119 general practices and 388 individual general practitioners, with three-quarters of participants referred directly by their family doctor. The mean age was seventy-nine point five years. The mean baseline score on the Alzheimer's Disease Assessment Scale cognitive subscale, a seventy-point scale where higher scores mean worse cognition, was twenty-two point seven. The elegance of the design is worth pausing on. The same one hundred seventy-six people answered both questions simultaneously. The drug comparison was double-blind. The follow-up comparison couldn't be blinded — you can't hide from a participant whether researchers are visiting every two months or just once — but it was randomised, which is the most important aspect.

After six months, more attention had produced a measurable cognitive benefit. In the primary intention-to-treat analysis, adjusting for baseline score, participants in the intensive follow-up group scored about two points better on the Alzheimer's Disease Assessment Scale cognitive subscale than those in the minimal group. The adjusted mean difference was minus two point zero two, with a ninety-five percent confidence interval ranging from minus three point nine one to minus zero point twelve, and a p-value of zero point zero three seven. The same direction held in the per-protocol analysis, where the difference was slightly larger at minus two point three eight. Two points is modest. McCarney and colleagues say so plainly: "This may not be regarded as particularly clinically significant." But they immediately note it is of similar magnitude to effects reported in randomised controlled trials of cholinesterase inhibitors, the standard drug class for Alzheimer's disease. The observation alone, in other words, produced a cognitive shift comparable to what some treatments claim. Then came the twist. Participant-rated quality of life — measured with the Quality of Life in Alzheimer's Disease, a thirteen-item scale where higher scores mean better quality of life — told a different story entirely. It favoured the minimal follow-up group.

The adjusted mean difference was minus one point three eight, with a ninety-five percent confidence interval of minus two point six four to minus zero point twelve, and a p-value of zero point zero three two. Patients who were seen less frequently felt slightly better about their lives. Carer-rated quality of life showed no significant difference either way. So the trial produced a genuine paradox. More intensive follow-up improved measured cognition but slightly worsened how patients rated their own well-being. The cognitive benefit and the quality-of-life cost pointed in opposite directions, and the data alone can't fully resolve why. McCarney and colleagues work through the candidates honestly. The most obvious explanation for the cognitive finding is a practice or learning effect — patients who take the Alzheimer's Disease Assessment Scale cognitive subscale four times may simply get better at the test. The researchers tried to limit this by rotating the word lists used in the memory items.

They also note the test-retest coefficient for the Alzheimer's Disease Assessment Scale cognitive subscale over six weeks is zero point nine one, which they cite as evidence against a large learning effect. But they admit they cannot rule it out, particularly in a dementia population where test familiarity and genuine cognitive change are hard to separate. A second candidate is social contact itself: more frequent visits from researchers may improve mood, motivation, or general engagement in ways that show up on cognitive tasks. A third is what the paper calls better recognition of needs — more contact may help participants or families identify and address problems they would otherwise miss. The quality-of-life finding points toward a different mechanism. More frequent reminders of being in a dementia study may increase awareness of the diagnosis and disability, lowering self-rated well-being even as it improves cognitive performance. The burden of repeated assessments — each taking between ninety minutes and two hours — may also weigh on participants in the intensive group in ways that do not register on a cognitive scale but do show up in how someone answers questions about their own life. Notably, carer-rated quality of life was unaffected, which suggests that the effect is happening in how patients perceive and report their own experience, not in anything their carers observe from the outside.

The authors are candid about one more complication: adherence to the protocol differed between groups. Nearly thirty percent of participants in the intensive group were not followed exactly according to protocol, compared to about fourteen percent in the minimal group. That likely means the true effect of follow-up intensity is underestimated, not inflated. What the DIGGER trial ultimately demonstrates has reach well beyond this particular study of dementia and Ginkgo biloba. If two points of cognitive improvement can emerge from observation alone — from the act of being followed closely, assessed repeatedly, and contacted by researchers — then the treatment effects reported in rigorous clinical trials may systematically overestimate what happens in routine clinical care. Most trials cannot isolate this component, because intensive monitoring typically applies equally to treatment and placebo arms. It gets subtracted out of the drug comparison but inflates both groups relative to what ordinary patients experience when they see their doctor twice a year. As McCarney and colleagues put it, such an effect "may result in an inflated estimate of effect size in routine clinical settings by over-estimating response in both groups."

They raise a provocative corollary. If participation in research itself produces benefit, one response would be to deliver routine care under conditions more similar to trials — closer monitoring and more frequent contact. The authors acknowledge this would have resource implications, then leave it there. But the logic holds: if the observation is therapeutic, the observation has value. The paper also notes that people with dementia may actually show a smaller Hawthorne Effect than cognitively intact populations, because their capacity for learning and social response may be reduced. The two-point Alzheimer's Disease Assessment Scale cognitive subscale shift observed here could be a floor, not a ceiling, for what this phenomenon does in other trials. Before this study, the Hawthorne Effect in clinical research was exactly what the original Hawthorne research had always been — a compelling story with contested evidence behind it. McCarney and colleagues ran the first randomised controlled trial designed to quantify it and found a small but real signal. Two points on a cognitive scale in community-dwelling patients with mild-to-moderate dementia. Modest, as the authors say. But it existed, it was measurable, and it showed up in both the cognitive data and the quality-of-life data, pointing in different directions at once. The effect that spent decades as a cautionary footnote now has a number.

Whether it comes from observation, practice, social contact, or something we haven't named yet — the paper is clear that it "remains intriguing and clinically relevant, whatever the cause." This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Medicine