Instagram photos reveal predictive markers of depression

Andrew Reece, Christopher M. DanforthView original
OverviewBalancedalloy voice
Depression is underdiagnosed, and the numbers are uncomfortable. Meta-analytic data shows that unassisted general practitioners correctly identify depressed patients only about fifty-one percent of the time, and on average, more than half of their depression diagnoses are false positives. That's the clinical reality Andrew Reece and Christopher Danforth set out to challenge. Their tool of choice is Instagram. The premise sounds almost too simple. People post photographs constantly, and those photos carry information about color, social context, mood, and activity levels. If those signals are consistent enough, a machine might read them better than a busy doctor working from a brief appointment. To test this, Reece and Danforth recruited one hundred sixty-six participants through Amazon's Mechanical Turk, running separate surveys for people with a history of depression and healthy controls. Depressed participants completed the Center for Epidemiologic Studies Depression Scale, a standard screening questionnaire, provided demographic and clinical history, and crucially, shared their Instagram usernames. After obtaining consent, the researchers collected each participant's entire posting history. The final dataset consisted of forty-three thousand nine hundred fifty photographs. Privacy was taken seriously. Participants were warned that strict anonymity was difficult to guarantee given that usernames and personal photos can be identifying. No personal identifiers, including usernames, were published. Participants could complete the survey only once, and access was restricted to experienced Mechanical Turk workers with high approval ratings in the United States. Now, what do you actually measure in a photograph? Reece and Danforth extracted features a human might not consciously notice. From each image, they computed pixel-level averages for hue, saturation, and value, according to the HSV color system. Hue describes where a color falls on the light spectrum, with lower values being redder and higher values being bluer. Saturation is vividness; low saturation makes an image look gray and washed out. Value is brightness; lower value means a darker image. Beyond color, they ran algorithmic face detection, recording whether a photo contained any human face and how many. They logged engagement metrics such as likes and comments per post and measured posting frequency. They also recorded filter use, not just whether a filter was applied, but which one. The paper notes that a computer can analyze the average saturation value of a million pixels across thousands of photos. That's the scale advantage. What they found consistently was that photographs posted by people with depression looked different in measurable ways: bluer, darker, and grayer. Higher hue, lower brightness, and lower saturation. Depressed users were less likely to apply any filter at all, and when they did, they disproportionately favored Inkwell, a black-and-white filter. Healthy controls disproportionately favored Valencia, which warms and lightens images. The face-detection results were more nuanced; depressed participants were more likely to post photos containing at least one face, but those photos contained fewer faces on average. Reece and Danforth read this as potentially consistent with smaller or more private social contexts. Comment and like patterns diverged too; more comments were associated with depression, while more likes predicted the healthy group. These features fed into two models. Bayesian logistic regression, a statistical approach that estimates the strength and direction of each predictor and compares model fits, identified which individual features mattered. A Random Forest classifier, an ensemble method that builds many decision trees and combines their outputs, was used to evaluate overall predictive performance. Here's where the comparison with human judgment becomes one of the paper's most striking findings. A separate group of Mechanical Turk raters evaluated a subset of photos on scales of happiness, sadness, likability, and interestingness—exactly the kind of intuitive impressions a person would form glancing at someone's feed. Human ratings of sadness and happiness were significant predictors of depression, but the human ratings model was weaker overall than the machine model. Moreover, human ratings were almost entirely uncorrelated with the computational features. The only notable exception was a modest positive relationship between how happy raters found a photo and the presence of faces. In other words, what a human notices in a photo and what the algorithm extracts are largely different things. The machine is picking up on signals that bypass conscious perception. The all-data classifier, trained on participants' full posting histories, achieved a recall of seventy percent, meaning it correctly identified about seventy percent of depressed individuals. Precision was sixty percent, and the F1 score was sixty-five percent. Compared to the general practitioner benchmark from Mitchell and colleagues, the all-data model outperformed in recall, precision, and F1. The one area where general practitioners held an advantage was specificity, correctly identifying healthy individuals, where the benchmark was eighty-one percent against the model's forty-eight percent. So the machine identifies more true cases but also generates more false positives. That's a meaningful trade-off, and Reece and Danforth don't hide it. But the pre-diagnosis result is where the research opens up into something genuinely significant. They trained a separate classifier using only posts made before depressed participants received their first clinical diagnosis. This model never saw any data from after someone entered the healthcare system, yet it still worked. The pre-diagnosis model was conservative; recall dropped to thirty-two percent, meaning it caught about a third of true cases. But when it did predict depression, it was more often correct; precision was fifty-four percent, and specificity rose to eighty-three percent, actually exceeding the general practitioner benchmark on that metric. The behavioral signal in these photographs predates formal clinical recognition. The photos were already changing—getting darker, grayer, and more filtered toward black-and-white—before a doctor ever made a diagnosis. The authors were careful to avoid a methodological trap here; they explicitly did not apply the larger all-data model to pre-diagnosis observations, because that would artificially inflate accuracy. The pre-diagnosis performance is likely an underestimate of what a properly designed prospective system could achieve. The limitations are real, and Reece and Danforth name them directly. The sample comes from Mechanical Turk workers, who may not represent the broader population of people with depression. The depression label is non-specific; it doesn't distinguish subtypes or account for comorbidities. Additionally, a portion of survey participants refused to share their Instagram data even after privacy assurances, which introduces a selection effect that's hard to fully characterize. These aren't fatal flaws, but they're constraints that any clinical application would have to take seriously. The suggested path forward is to add text—captions, comments, and hashtags—the linguistic layer that accompanies every post. Prior work in health detection from social media has leaned heavily on text, and Reece and Danforth argue that combining visual and textual features in a multimodal model could outperform either alone. The images are already carrying a signal. The words might carry a different one. Together, they might form something more complete. The privacy tension is real and worth considering. The photographs in this study were posted publicly, or at least shared with the platform. However, using them to infer mental health status feels qualitatively different from browsing someone's feed. Reece and Danforth flag this directly, noting that socio-technical trends shift over time, and that any deployed system would require frequent recalibration. Consent, transparency, and ongoing validation aren't optional features of a clinical screening tool; they're the foundation. Depression affects hundreds of millions of people globally. Diagnosis is often delayed by years. The gap between onset and treatment is where lives deteriorate, relationships break down, and interventions become harder. A passive, low-cost screening signal that operates at scale—not replacing clinical judgment, but flagging people who might benefit from a conversation—could matter enormously in settings where mental health services are scarce. What Reece and Danforth demonstrated is that the signal exists, that it predates diagnosis, and that a machine can find it in places a busy clinician would never think to look: the color palette of someone's photos, the filters they choose, the faces they share, and the quiet visual texture of a life shared online. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Depression is underdiagnosed, and the numbers are uncomfortable. Meta-analytic data shows that unassisted general practitioners correctly identify depressed patients only about fifty-one percent of the time, and on average, more than half of their depression diagnoses are false positives. That's the clinical reality Andrew Reece and Christopher Danforth set out to challenge. Their tool of choice is Instagram.

The premise sounds almost too simple. People post photographs constantly, and those photos carry information about color, social context, mood, and activity levels. If those signals are consistent enough, a machine might read them better than a busy doctor working from a brief appointment.

To test this, Reece and Danforth recruited one hundred sixty-six participants through Amazon's Mechanical Turk, running separate surveys for people with a history of depression and healthy controls. Depressed participants completed the Center for Epidemiologic Studies Depression Scale, a standard screening questionnaire, provided demographic and clinical history, and crucially, shared their Instagram usernames. After obtaining consent, the researchers collected each participant's entire posting history.

The final dataset consisted of forty-three thousand nine hundred fifty photographs.

Privacy was taken seriously. Participants were warned that strict anonymity was difficult to guarantee given that usernames and personal photos can be identifying. No personal identifiers, including usernames, were published.

Participants could complete the survey only once, and access was restricted to experienced Mechanical Turk workers with high approval ratings in the United States.

Now, what do you actually measure in a photograph? Reece and Danforth extracted features a human might not consciously notice. From each image, they computed pixel-level averages for hue, saturation, and value, according to the HSV color system.

Hue describes where a color falls on the light spectrum, with lower values being redder and higher values being bluer. Saturation is vividness; low saturation makes an image look gray and washed out. Value is brightness; lower value means a darker image.

Beyond color, they ran algorithmic face detection, recording whether a photo contained any human face and how many. They logged engagement metrics such as likes and comments per post and measured posting frequency. They also recorded filter use, not just whether a filter was applied, but which one.

The paper notes that a computer can analyze the average saturation value of a million pixels across thousands of photos. That's the scale advantage. What they found consistently was that photographs posted by people with depression looked different in measurable ways: bluer, darker, and grayer.

Higher hue, lower brightness, and lower saturation. Depressed users were less likely to apply any filter at all, and when they did, they disproportionately favored Inkwell, a black-and-white filter. Healthy controls disproportionately favored Valencia, which warms and lightens images.

The face-detection results were more nuanced; depressed participants were more likely to post photos containing at least one face, but those photos contained fewer faces on average. Reece and Danforth read this as potentially consistent with smaller or more private social contexts. Comment and like patterns diverged too; more comments were associated with depression, while more likes predicted the healthy group.

These features fed into two models. Bayesian logistic regression, a statistical approach that estimates the strength and direction of each predictor and compares model fits, identified which individual features mattered. A Random Forest classifier, an ensemble method that builds many decision trees and combines their outputs, was used to evaluate overall predictive performance.

Here's where the comparison with human judgment becomes one of the paper's most striking findings. A separate group of Mechanical Turk raters evaluated a subset of photos on scales of happiness, sadness, likability, and interestingness—exactly the kind of intuitive impressions a person would form glancing at someone's feed. Human ratings of sadness and happiness were significant predictors of depression, but the human ratings model was weaker overall than the machine model.

Moreover, human ratings were almost entirely uncorrelated with the computational features. The only notable exception was a modest positive relationship between how happy raters found a photo and the presence of faces. In other words, what a human notices in a photo and what the algorithm extracts are largely different things. The machine is picking up on signals that bypass conscious perception.

The all-data classifier, trained on participants' full posting histories, achieved a recall of seventy percent, meaning it correctly identified about seventy percent of depressed individuals. Precision was sixty percent, and the F1 score was sixty-five percent. Compared to the general practitioner benchmark from Mitchell and colleagues, the all-data model outperformed in recall, precision, and F1.

The one area where general practitioners held an advantage was specificity, correctly identifying healthy individuals, where the benchmark was eighty-one percent against the model's forty-eight percent. So the machine identifies more true cases but also generates more false positives. That's a meaningful trade-off, and Reece and Danforth don't hide it.

But the pre-diagnosis result is where the research opens up into something genuinely significant. They trained a separate classifier using only posts made before depressed participants received their first clinical diagnosis. This model never saw any data from after someone entered the healthcare system, yet it still worked.

The pre-diagnosis model was conservative; recall dropped to thirty-two percent, meaning it caught about a third of true cases. But when it did predict depression, it was more often correct; precision was fifty-four percent, and specificity rose to eighty-three percent, actually exceeding the general practitioner benchmark on that metric. The behavioral signal in these photographs predates formal clinical recognition.

The photos were already changing—getting darker, grayer, and more filtered toward black-and-white—before a doctor ever made a diagnosis. The authors were careful to avoid a methodological trap here; they explicitly did not apply the larger all-data model to pre-diagnosis observations, because that would artificially inflate accuracy. The pre-diagnosis performance is likely an underestimate of what a properly designed prospective system could achieve.

The limitations are real, and Reece and Danforth name them directly. The sample comes from Mechanical Turk workers, who may not represent the broader population of people with depression. The depression label is non-specific; it doesn't distinguish subtypes or account for comorbidities.

Additionally, a portion of survey participants refused to share their Instagram data even after privacy assurances, which introduces a selection effect that's hard to fully characterize. These aren't fatal flaws, but they're constraints that any clinical application would have to take seriously.

The suggested path forward is to add text—captions, comments, and hashtags—the linguistic layer that accompanies every post. Prior work in health detection from social media has leaned heavily on text, and Reece and Danforth argue that combining visual and textual features in a multimodal model could outperform either alone. The images are already carrying a signal.

The words might carry a different one. Together, they might form something more complete.

The privacy tension is real and worth considering. The photographs in this study were posted publicly, or at least shared with the platform. However, using them to infer mental health status feels qualitatively different from browsing someone's feed.

Reece and Danforth flag this directly, noting that socio-technical trends shift over time, and that any deployed system would require frequent recalibration. Consent, transparency, and ongoing validation aren't optional features of a clinical screening tool; they're the foundation.

Depression affects hundreds of millions of people globally. Diagnosis is often delayed by years. The gap between onset and treatment is where lives deteriorate, relationships break down, and interventions become harder.

A passive, low-cost screening signal that operates at scale—not replacing clinical judgment, but flagging people who might benefit from a conversation—could matter enormously in settings where mental health services are scarce. What Reece and Danforth demonstrated is that the signal exists, that it predates diagnosis, and that a machine can find it in places a busy clinician would never think to look: the color palette of someone's photos, the filters they choose, the faces they share, and the quiet visual texture of a life shared online.

This lecture was created by ennepō.

Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field.

Read when you can. Listen when you want to.

More in Psychology