Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot)A Randomized Controlled Trial
Let’s start with the why. College is supposed to be a launchpad for possibility, but for a lot of students, it’s also the first time anxiety and depression really bite down. Most mental health diagnoses start young — roughly three out of four before age twenty-four — and on campus, more than half of students report symptoms in the prior year so severe they interfered with daily life.
Even when people want help, they don’t always get it. Stigma still pushes students away from care, and the result is a stubborn gap where as many as three in four who need services never access them. That isn’t an abstract statistic.
It’s missed classes, lost sleep, and a mind that feels like it’s constantly flickering.
Digital tools were supposed to help close that gap. Interest is there — about seven in ten people say they would use a phone app to monitor or manage their mental health — and internet programs for anxiety and depression can lead to outcomes comparable to therapist-delivered cognitive behavioral therapy, or CBT. But in the wild, the reality has been messy.
People don’t stick with these tools. A typical program sees barely half its users complete even minimal content, and a review a decade ago could count thousands of abstracts but only a handful of apps that had been tested in randomized trials, most of them not commercially available. The problem isn’t just content.
It’s the feel. We show up for people, not portals, and a flat, self-guided program rarely creates the sense of being met that keeps us engaged.
That’s the wager behind Woebot. Take the core moves of CBT — noticing thoughts, testing beliefs, taking small, values-consistent actions — and wrap them in a conversation. Not a human behind a screen, but a fully automated, text-based agent designed to mirror the rhythm of therapy: brief check-ins, a bit of psychoeducation, a quick exercise, and then back you go into your day.
The design team leaned hard into evidence-based app heuristics. They targeted both anxiety and low mood, tailored content to what users reported, prompted people to capture thoughts and feelings, linked activities to the problems they named, encouraged non-screen actions, used light gamification and reminders to keep things moving, and kept the interface simple with crisis-support links a tap away. In other words, not a chatbot that happens to mention CBT, but a CBT program that happens to chat.
Fitzpatrick and colleagues put that idea into a small, pragmatic test: a two-week randomized controlled trial among young adults with depressive and anxious symptoms. Recruitment was simple — social media flyers across a U.S. university community — with basic checks to confirm eligibility and screen out bots. After consenting, participants were randomized by algorithm to one of two things.
If they drew Woebot, they got a direct link to the agent in an instant messenger. If they drew control, they got the National Institute of Mental Health’s ebook on depression in college students. Seventy people made it through screening and randomization — thirty-four to Woebot and thirty-six to information control — after removing a tranche of automated sign-ups.
The groups looked broadly similar at baseline. Depressive symptoms on the Patient Health Questionnaire-9, or PHQ-9, hovered around the mid-teens, roughly fourteen in the Woebot arm and thirteen in control. Anxiety on the Generalized Anxiety Disorder scale clustered near the high-teens to twenty, about eighteen versus nineteen, and nearly three-quarters met the “severe” threshold for anxiety.
Ages sat in the early twenties. Allocation was concealed by the algorithm, although staff knew which condition people were in. Analytically, the team planned an intention-to-treat analysis using analysis of covariance to adjust for baseline scores, handled missing data with multiple imputation, and had powered the study so that seventy participants would give about eighty percent power to detect a moderate effect on depression. It’s a tight design for a feasibility window.
Before we talk outcomes, it’s worth asking: did people even use it? Yes, and at a near-daily clip. Over two weeks, Woebot users averaged about twelve check-ins, with a median of twelve and a range from eight to eighteen.
Most of those happened on unique days, which means the bot didn’t just get sampled; it became a brief daily practice. Conversations were quick — anywhere from ninety seconds to about ten minutes — depending on whether the moment called for a small lesson or a skill exercise. Engagement in the ebook group couldn’t be tracked directly; the researchers only had participants’ comments to go on, and thirteen people in that arm described reading the material at least once.
That asymmetry matters when we think about dose, but it also tells us something pragmatic: a conversational nudge can pull you in.
On the primary outcome of depression at about two weeks, Woebot did better than information alone. The adjusted mean PHQ-9 score in the Woebot group landed at 11.14, compared with 13.67 in control. Statistically, that difference cleared the bar with an F of 6.03 and a p-value of 0.017, and it mapped to a moderate between-group effect size of 0.44.
The signal held up after correction for multiple comparisons. In a short window, with no human in the loop, a text agent nudged depressive symptoms down more than an ebook did.
Anxiety painted a more complicated picture. Across completers in both arms, anxiety eased over time with a modest within-person effect — about a third of a standard deviation — but there was no reliable difference between Woebot and the information control at that two-week mark. In plain terms, people’s anxiety improved a bit no matter which path they took, and the between-group gap never quite opened.
Affect, measured by positive and negative mood scales, showed no differences between groups either.
Attrition tells its own story about acceptability. Overall follow-up at two weeks was strong at eighty-three percent, but the split by arm was striking: thirty-one percent of the control group fell away versus nine percent in the Woebot group, a difference that cleared significance with a p-value of 0.023. When nearly a third of one group drifts off while the other mostly sticks around, you’re seeing not just an outcome signal, but an engagement signal. People came back for the conversation.
Satisfaction echoed that. On a five-point scale, overall satisfaction was higher among Woebot users — 4.3 compared with 3.4 in the ebook group, with a p-value below 0.001 — and they were more likely to say the experience sharpened their emotional awareness, 3.3 versus 2.7 on average. The head-turner is this: every single Woebot participant endorsed learning something from the experience, compared with about three out of four in the information arm.
That isn’t a knock on the quality of the ebook. It’s a hint that interaction — being asked a question, being remembered — changes how we take in the very same ideas.
If you listen to what participants actually said, you can hear the contours of that relationship. On the “best thing” list, accountability cropped up again and again — nine people named the daily check-ins as the engine that kept them honest — and seven called out Woebot’s empathy or personality. Learning was another strong theme, with a dozen comments about new insights into emotions or thinking patterns.
On the flip side, when the process broke — a tone that missed, a response that felt canned — people bristled; fifteen tagged those process violations as the worst part, with technical glitches and content misses a distant second. And there’s this wonderfully human detail: several students described Woebot as “he,” a “friend,” even a “fun little dude.” That doesn’t mean the bot is a therapist. It means that, for two weeks, a scripted exchange could approximate enough of a therapeutic bond to matter.
Under the hood, that bond was built on structure. Woebot’s engine blended short psychoeducation with guided exercises, took mood and context snapshots during check-ins, and nudged people toward small actions in their real lives. Safety scaffolding was there — crisis resources and helpline numbers when risk was elevated — but no human clinician was steering individual chats.
Ethically, the study cleared institutional review with brief, checkbox consent suitable for an online trial. Participants were compensated — ten dollars per completed assessment, twenty if they did both — and usage data were de-identified to protect privacy, which had a trade-off: the team couldn’t link intensity of use to individual outcomes. One more important disclosure: a senior author was a founder of the company behind Woebot and covered participant incentives.
That kind of conflict doesn’t negate the results, but it deserves to be on the table.
There’s also a curious exploratory finding inside the depression scores. When Fitzpatrick’s team looked item by item within the Woebot group, the biggest shift wasn’t on sadness or sleep. It was on psychomotor symptoms — feeling slowed down or restless — with a very large within-person change, a d of 2.09.
Appetite moved next, at 0.65, followed by smaller but notable changes in anhedonia — that “little interest or pleasure” item — at 0.44, and self-worth at 0.40. It’s one study, short and small, so we shouldn’t over-interpret item patterns. But it suggests the agent’s blend of behavioral tips and cognitive reframes might first land on energy, activity, and how people talk to themselves.
Every feasibility trial has guardrails, and this one is no exception. It’s two weeks long, with seventy randomized and fifty-eight providing data at follow-up. No long-term follow-up, no way to estimate durability.
Engagement in the control arm couldn’t be tracked objectively, which muddies dose–response comparisons. Analyses were rigorous for what they were — intention-to-treat with baseline adjustment and multiple imputation — but the sample was too small to test mediation, so we can’t say whether daily check-ins drove symptom change. The population was a college sample from a specific region, with no formal assessment of digital access or equity, so generalizability is an open question.
And remember the conflict of interest. All of this tempers the headline without erasing it.
So what do we carry forward? First, the idea that conversation itself may be the engagement technology. Prior app trials stumbled not because CBT stopped working, but because flat delivery couldn’t hold attention.
Here, near-daily check-ins and higher satisfaction suggest that a carefully designed agent can evoke enough of a social feel to keep people coming back, at least for a short run. Second, the outcome split matters. Depression moved with Woebot beyond information alone; anxiety did not, at least not between groups in two weeks.
That could be timing, content, or simply power, but it sets a clear agenda for follow-up.
And that agenda is not mysterious. Fitzpatrick and colleagues are explicit: larger, longer trials with active, blinded controls and real follow-up are needed to see what lasts and why. Bring in objective engagement metrics for all arms.
Test mechanisms directly — does accountability mediate change, or is it the psychoeducation, or the cognitive reappraisal practice? And run these trials in more diverse settings, from community colleges to non-student young adults, to see where a chatbot slots into the real ecosystem of care.
One cautious optimism is allowed. If a fully automated agent can deliver a modest, meaningful lift in depressive symptoms over two weeks, with high engagement and no human clinician in the loop, that’s a potentially scalable piece of the puzzle. Not a replacement for therapy.
Not a cure-all. A small, daily companion that asks good questions, remembers your answers, and nudges you toward the next right thing. In a world where the bottleneck is often getting through the door, that might be the difference between starting and stalling.
Related lectures
- Insomnia and the risk of depression: a meta-analysis of prospective cohort studies
- Mirror-Induced Behavior in the Magpie (Pica pica): Evidence of Self-Recognition
- The cross-national epidemiology of social anxiety disorder: Data from the World Mental Health Survey Initiative
- The Natural Statistics of Audiovisual Speech
- The Small World of Psychopathology
- Health-related quality of life in parents of school-age children with Asperger syndrome or high-functioning autism