Mobile Phone Sensor Correlates of Depressive Symptom Severity in Daily-Life BehaviorAn Exploratory Study
Imagine if your phone could help flag when your mood is sliding long before you say it out loud. Not by asking you questions, but just by noticing how you move through the world and how often you tap the screen. That’s the promise behind a small but provocative study by Saeb and colleagues, who set out to see whether passive smartphone data could trace the contours of depressive symptoms.
The stakes are real: major depression is common, with annual prevalence often landing between about seven and ten percent and lifetime risk around seventeen percent. Below the diagnostic threshold, many more people feel the drag of low mood and low energy. Yet in everyday care, we miss a lot of it. Phones, carried in pockets all day, might help us notice.
Here’s the question they asked, very simply: can everyday geographic movement and phone use carry a signature of depressive symptoms as measured by the Patient Health Questionnaire-9, or PHQ-9, a standard nine-item questionnaire? To test it, they used a custom Android app called Purple Robot and followed people for two weeks. Forty adults enrolled.
By the end, usable sensor data came back from twenty-eight of them. At the start, everyone filled out the PHQ-9, which gave the ground truth. Half of the analyzable group had scores below five, the “no symptoms” range, and half had scores of five or more, from mild up to more severe.
The sample was young on average—just under twenty-nine—and skewed female. Small, yes, but enough to see whether the signals were even there.
Let’s talk about what the phone actually measured. Purple Robot sampled GPS location every five minutes. It also watched for screen on and off events to estimate how often, and for how long, the phone was used.
Very brief screen flickers not initiated by the person—anything under thirty seconds—were tossed out so they wouldn’t inflate the usage signal. The app buffered data on the device when the network was spotty, then uploaded later using a store-and-forward design. Identifiers were anonymized with hashing and the payload was encrypted before transmission.
Not glamorous details, but they matter when you’re asking people to let a sensor trail them for weeks.
From there, the raw stream became behavior. The team split GPS points into stationary versus in-motion states using a simple rule of thumb: if the inferred speed was under about one kilometer per hour, you were considered stationary. Only those stationary points were clustered—using K-means—into the places you spend most of your time: home, work, the gym, a favorite café.
They expanded the number of clusters until the farthest point in each cluster was within roughly five hundred meters of its center. “Home” wasn’t hard-coded; it was defined as one of the top three most visited clusters and also the one most visited between midnight and six a.m.
A few features came out of this process that you can probably feel intuitively. Home stay is the percent of time you spend at the home cluster. Normalized entropy refers to how evenly your time is spread across your usual locations, scaled so a person with more places isn’t penalized for having more categories.
If you spend all day in one spot, entropy drops; if you make the rounds among many places, entropy rises. Location variance measures how much you move around geographically regardless of which places those are. And circadian movement provides a measure of how regular your twenty-four-hour movement rhythm is.
To compute that last one, they used the Lomb–Scargle method—think of it as a way to pull out rhythmic energy from unevenly sampled data—and looked specifically at the energy around a twenty-four-hour cycle, targeting periods between about twenty-three and a half and twenty-four and a half hours. Then they log-transformed it because this kind of measure tends to be skewed. On the phone side, the features were simpler: average daily usage duration and average daily usage frequency.
So, did these behavioral snapshots line up with symptoms? They did, and in ways that make sense for how depression can shrink a person’s world. Among the GPS measures, circadian movement had a strong negative relationship with PHQ-9 scores: the less regular your day-night movement rhythm, the higher your score tended to be.
The correlation was about minus zero point sixty-three and statistically convincing. Normalized entropy, that “how evenly do you spread yourself across your places” metric, was also negative at roughly minus zero point fifty-eight. Same story for location variance, also around minus zero point fifty-eight.
Taken together, people with higher symptom scores were moving less, visiting fewer places, and following less regular daily movement rhythms. It fits the lived experience: depression often means staying home, canceling plans, and losing the clockwork that keeps days distinct.
Phone use told a different, but complementary, story. More time using the phone and more frequent screen checks went with higher PHQ-9 scores. Usage duration correlated around zero point fifty-four; usage frequency around zero point fifty-two.
In other words, people reporting more depressive symptoms were picking up the phone more often and for longer. Those two usage metrics were themselves very tightly linked, with a correlation close to zero point nine. That interdependence would matter later when combinations of features were tested.
Correlations are one thing. The team then pushed toward prediction. First, they asked a basic classification question: could a simple model separate people with any depressive symptoms—PHQ-9 of five or more—from those with none?
Using just one feature at a time, the standout was normalized entropy. On its own, it achieved about eighty-six point five percent accuracy. Sensitivity—how often the model correctly flagged someone in the symptoms group—was about eighty-eight percent.
Specificity—how often it correctly left people in the no-symptoms group alone—was around eighty-five percent. That’s a strong separation from a behavioral measure you can collect without asking a single question.
Other single GPS features did fine but trailed that front-runner. Circadian movement hit the high seventies for accuracy. Location variance and home stay sat in the mid-seventies.
Plain entropy—without the normalization step—dropped below seventy percent. Naive mobility counts did poorly: total distance traveled nudged the mid-fifties, and the raw number of location clusters hovered in the low forties. Interestingly, throwing all GPS features together did not beat the single best one.
The full bundle landed just under seventy-nine percent accuracy, with sensitivity in the mid-eighties and specificity in the mid-seventies. That pattern—one clean feature outperforming a kitchen sink—will be familiar to anyone who has wrestled with correlated predictors on small datasets.
On the regression side, the goal was to estimate each person’s PHQ-9 score rather than just group them. Here the error metric was the normalized root mean square deviation, or NRMSD, which simply means “how far off are we, as a fraction of the total score range we observed?” They normalized by the observed range of zero to seventeen. Using single GPS features, average errors clustered in a narrow band.
Circadian movement yielded an NRMSD around zero point two two two, normalized entropy was about zero point two three five, location variance sat near zero point two two nine, and home stay was closer to zero point two five three. Combining all GPS features did not improve that number; the bundle landed around zero point two five one. So we’re talking errors on the order of twenty to twenty-five percent of the full PHQ-9 span. Not diagnostic precision, but not noise either.
What about the phone-only models? On classification, usage duration alone reached roughly seventy-four percent accuracy. Its sensitivity was lower—mid-sixties—while specificity climbed into the mid-eighties.
Usage frequency alone was weaker, with accuracy around sixty-nine percent. Putting both usage features together actually made things worse, slipping toward sixty-six percent accuracy. On the regression task, NRMSD hovered in a similar range: roughly zero point two six eight for usage duration, about zero point two four nine for usage frequency, and around zero point two seven three when both were combined.
Again, correlated inputs and a small sample can make combinations underperform the best-behaved single feature.
A quick note on how these models were built, because that’s where statistical rigor lives or dies. For classification, they used logistic regression, which you can think of as taking a weighted sum of features—start with a constant, add each feature times its weight—and then passing that sum through an S-shaped curve that turns any number into a probability between zero and one. They used a probability threshold of zero point five to say “symptoms” or “no symptoms.” For regression, it was the classic least squares approach: choose the feature weights that make the predicted PHQ-9 scores land as close as possible to the real ones.
Because the sample was small and several features moved together, they used elastic-net regularization when multiple predictors were in play. Practically, that means adding a penalty to the cost function that grows with two things at once: the sum of the absolute values of the weights and the sum of their squares. The balance between those two penalties was tuned by cross-validation.
They also stress-tested their findings by bootstrapping—resampling features a thousand times—and by evaluating performance with leave-one-participant-out cross-validation so that each person served as a held-out test at least once.
Let’s ground this back in the people behind the numbers. Of the forty who signed up, twelve couldn’t be used because of sensor dropouts, charging gaps, or server hiccups. Among the remaining twenty-eight, PHQ-9 scores ranged from zero to seventeen, averaging around five and a half, and the groups split evenly into “none” versus “some” symptoms.
For the GPS analyses, the usable subset was smaller still—nine per group—because of missing location data. For phone use, ten in the symptoms group and eleven in the no-symptoms group contributed enough events. Those are thin slices.
And the features were not independent: normalized entropy, location variance, and home stay share information, and the two usage measures were highly redundant. All of this makes the results encouraging, but also tentative.
The caveats flow from that reality. Association does not prove causation. People who move less may score higher on the PHQ-9, but the phone can’t tell you why.
The sample is young and self-selected through an online ad, not a random draw from a clinic population. Multiple comparisons were explored without formal corrections. And while the models were carefully cross-validated, the gap between performance in a pilot and performance in a larger, more diverse cohort can be wide.
Saeb and colleagues are explicit about this: treat these as preliminary signals that need replication, not as clinical tools ready for deployment.
Still, the throughline is clear. Two weeks of passive sensing, with nothing more exotic than GPS and screen events, captured enough about daily rhythms and movement to relate meaningfully to depressive symptom burden. A single, interpretable feature—the evenness of time spent across your usual places—could separate people with and without any symptoms with roughly eighty-six percent accuracy.
Simple rhythm measures of a twenty-four-hour day, derived from irregular samples using a spectral method built for astronomy, were among the best predictors of score. And phone engagement itself—how often and how long—carried signal too, albeit less cleanly than geography.
If you’re picturing where this could go, keep two things in mind. First, privacy and consent are not footnotes. The team anonymized and encrypted data end-to-end, and any real-world use would have to double down on transparency and control.
Second, the jump from "we can detect a pattern" to "we should act on it" is a major one. The future-facing idea is almost irresistible: continuous, low-burden monitoring that notices when someone’s life is narrowing and nudges them—or their clinician—at the right moment. But the right next step isn’t to roll that out; it’s to replicate these findings in bigger, more varied samples, link them to clinical outcomes, and understand how to intervene ethically when the phone thinks you’re not doing well.
In the end, what Saeb and colleagues showed is both modest and exciting. Modest because it’s an exploratory study with all the limitations that come with that label. Exciting because a phone in your pocket, watching how you move and how often you wake its screen, can sketch an outline of your mental state—enough to be useful, if we sharpen the picture and use it with care.
Related lectures
- Insomnia and the risk of depression: a meta-analysis of prospective cohort studies
- Mirror-Induced Behavior in the Magpie (Pica pica): Evidence of Self-Recognition
- The cross-national epidemiology of social anxiety disorder: Data from the World Mental Health Survey Initiative
- The Natural Statistics of Audiovisual Speech
- The Small World of Psychopathology
- Health-related quality of life in parents of school-age children with Asperger syndrome or high-functioning autism