Depression sum-scores don’t add upwhy analyzing specific depression symptoms is essential
Let's start with a simple picture that turns out to be wrong. Most research and a lot of clinical practice treat depression like a single thing you can add up. Count the symptoms, cross a threshold, and you're "depressed." It's tidy.
It's also, as Eiko Fried and Randolph Nesse put it, a convenient fiction. When you actually look under the hood, the pieces don't move together the way that story suggests.
Take the official checklist. The Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition, lays out nine symptom criteria: low mood, loss of interest, changes in sleep, appetite or weight, psychomotor slowing or agitation, fatigue, guilt or worthlessness, trouble concentrating or indecision, and suicidal thoughts. You need five of these, and one of them has to be low mood or loss of interest.
On paper that sounds precise. In reality, three of those items each contain opposites—insomnia or hypersomnia, weight gain or loss, agitation or retardation—which means you can assemble roughly one thousand different combinations that all qualify as major depressive disorder. Many of those combinations share no symptoms at all.
So, two people can both cross the diagnostic line yet be living with very different problems.
Here's where the fiction bites. If you add those symptoms into a single sum score, you're assuming they're interchangeable indicators of one underlying condition. Fried and Nesse went looking for cracks in that assumption by asking a simple question: which symptoms actually track with the way people are impaired in daily life?
In a sample of three thousand seven hundred three outpatients with depression, they found striking differences. Sad mood carried a big share of the story, accounting for 20.9 percent of the explained variance in functional impairment. Hypersomnia, which is sleeping too much, barely registered at 0.9 percent.
That's not a rounding error. And the details matter across life domains: loss of interest hits social activities hardest, while fatigue tends to sink home management. Two people could score the same total on a scale and be worlds apart in how badly their lives are disrupted and in what ways.
Those differences aren't just cosmetic. They map onto biology. When researchers break depression down to individual symptoms, heritability is all over the place—from near zero up to about 35 percent.
Somatic changes like loss of appetite and loss of libido, and cognitions like guilt or hopelessness, show stronger genetic influence than more general negative affect or tearfulness. Specific genetic links show up too: middle insomnia on the Hamilton Rating Scale for Depression has been tied to a particular haplotype in TPH1, which is a gene in the serotonin pathway. Large twin studies have teased out three genetic factors that load differently on different symptoms, leading those authors to say the old idea of a single genetic liability for "depression" doesn't fit.
And in the high-stakes territory of suicide risk, work on SKA2 polymorphisms found that the genes didn't act in a vacuum—interactions with anxiety and stress explained a substantial part of who attempted or thought about suicide. If you're only staring at a total score, those signals blur.
Inflammation tells a similar story. Sometimes it's elevated in depression, but not for everyone. In fact, fewer than half of people with depression show heightened inflammatory markers, and those cytokines aren't specific to depression anyway.
But when you look at clusters of symptoms, a pattern emerges. Sleep disturbance and appetite or weight changes—the bodily sides of depression—tend to be more common in inflamed states. That pushes us toward endophenotypes, more fine-grained subprofiles that might actually map to mechanisms, rather than one-size-fits-all labels.
Treatment response lines up with that symptom-centric view. Uher and colleagues reported a split that's as practical as it is interesting: people with high baseline systemic inflammation did better on nortriptyline, which is a tricyclic antidepressant, while those with low inflammation improved more on escitalopram, which is a selective serotonin reuptake inhibitor. Sleep problems often undercut the effectiveness of any antidepressant; persistent insomnia roughly doubles the chance that someone stays depressed.
Anxiety symptoms drag down remission rates, and some cognitive-motivational symptoms—loss of interest, diminished activity, indecision—predict weaker response. Layer on the fact that antidepressant side effects overlap with symptoms on legacy scales, and you can see why a single severity number makes it hard to tell if a person is not improving or simply hit by side effects that increase their score.
Life events also don't push every lever equally. After a romantic breakup, you tend to see low mood and guilt surge. Long-term stress is more likely to bring on fatigue and hypersomnia.
Across studies, adverse events forecast increases in particular symptoms, not a generic and uniform worsening. That should change how we design studies and how we talk to patients—matching context to expected symptom trajectories, rather than treating all stressors as the same fuel for the same fire.
Sleep is a great case study in this symptom-specific logic. A meta-analysis parsing different sleep problems found that insomnia, parasomnias, and sleep-related breathing disorders are linked to higher suicidality across psychiatric conditions, while hypersomnia doesn't show the same association. In lab experiments where researchers kept people up, performance cratered.
Compared to well-rested controls, sleep-deprived participants were about 0.87 standard deviations worse on psychomotor tasks and 1.55 standard deviations worse on cognitive tasks. Mood fell a dramatic 3.16 standard deviations. Here's a vivid way to picture that: the median sleep-deprived person performed like someone at the ninth percentile of rested controls.
Other meta-analyses add that people with psychiatric diagnoses and sleep disturbances are about twice as likely to report suicidal behaviors as those without sleep problems. Sleep isn't one switch. Different sleep problems carry different risks.
So if the parts really do behave differently, we need methods that can see parts and the links between them in motion. Group averages can't do that on their own. Experience sampling methods—asking people about mood, activity, sleep, and symptoms multiple times a day—let us watch symptoms over time.
In one study, last night's sleep quality predicted how people felt the next day, but daytime mood didn't reliably predict how they'd sleep that night. That's a directional effect you'd miss if you only had before-and-after snapshots. Bringmann and colleagues tracked rumination and found huge variation in how sticky it was; for some people, today's rumination strongly predicted tomorrow's, while for others it didn't carry over.
Work on physical activity and depression has turned up similar heterogeneity in what causes what, with the direction of influence flipping across individuals. The lesson is not that there's no general pattern—it's that general patterns hide meaningful personal dynamics.
Measurement either helps you see that or it hides it. The field still leans hard on sum scores from older scales like the Hamilton Rating Scale for Depression and the Beck Depression Inventory, which blend together very different symptom domains and mix in somatic items that can be side effects of medication. Structured clinical interviews, like the Structured Clinical Interview for DSM Disorders, save time by skipping whole sections if a person doesn't meet core criteria.
That efficiency comes at a cost: you lose data on "non-core" symptoms that may be driving impairment or marking specific biology. Fried and Nesse argue for simple upgrades. Use multi-item measures for each symptom—suicidality, for instance, is assessed with six items in the Inventory of Depression and Anxiety Symptoms, which makes it more reliable.
Drop the skip logic when possible so you don't blind yourself to patterns outside the main criteria. And don't stop at DSM symptoms. Anxiety and anger are common in depression and linked to worse outcomes; the Symptoms of Depression Questionnaire includes them, and that kind of breadth pays off.
On the analysis side, there are tools built for this job. Item response theory can show you how each item behaves across the severity spectrum and whether a particular symptom functions differently in different groups—a phenomenon called differential item functioning. Structural equation modeling helps tease apart relationships among symptoms and latent traits, and it can model residual dependencies, the leftovers that ordinary factor models often ignore.
Fried and Nesse point to work showing that risk factors like neuroticism or past adverse events don't lift all symptoms equally; in one analysis that tested 25 predictors against nine symptoms, the effects were strikingly uneven. Attempts to carve out neat subtypes with latent class or factor-analytic techniques often fail to produce stable, clinically useful categories, even though the first factor in a scale typically explains a lot of variance. That's a hint: a general depression factor is real, but it isn't the whole story.
Network models go the next step and treat symptoms as nodes that can activate each other—insomnia raising fatigue, fatigue lowering activity, and lower activity darkening mood—potentially creating feedback loops that sustain illness. You don't have to buy the network view wholesale to see the value of modeling those links.
What does all this change in practice? For clinicians and researchers, it suggests a shift from "how depressed is this person?" to "which problems are present, how do they hang together, and which ones matter most for function and risk?" That means measuring insomnia and hypersomnia separately, not as a single sleep problem. It means tracking nightmares, psychomotor changes, fatigue, cognitive issues, anger, and anxiety alongside the classic nine symptoms.
It means collecting enough items per symptom to feel confident in your scores and using multiple scales to cover blind spots. It means analyzing symptoms on their own and together—mapping who is most impaired by what and who is likely to improve on which treatment. And when you run trials, report symptom profiles and consider stratifying by biomarkers when they're relevant.
Inflammation is a good start: people with higher baseline inflammation tend to show more somatic symptoms, and, as Uher's group showed, may respond better to nortriptyline than to escitalopram.
The stakes here are not just academic. Remember that sad mood's 20.9 percent share of explained impairment from that large outpatient study versus hypersomnia's 0.9 percent. That's a strong signal telling us where attention might buy the most benefit.
Remember that persistent insomnia doubles the chance of staying depressed, and that sleep loss can make an average person perform like a bottom decile control. Those are levers you can pull if you've measured them properly and taken them seriously.
There are limits. There isn't a single validated biomarker for depression, and heterogeneity is the rule, not the exception. Some scales have reliability issues, and not every lab has the bandwidth for dense, repeated measures.
But the path forward is unusually clear for a field that's been stuck. Measure richly at the symptom level. Use tools—item response theory, structural equation modeling, network approaches—that can detect unevenness and directionality.
Link symptoms to mechanisms where you can, and to function where you must. And then report it that way so other people can build on it.
If there's a unifying thought to leave with, it's this: progress accelerates when we respect the parts. Depression is not one thing, and treating it like one has muffled important signals in our data and in people's lives. The work that Fried and Nesse synthesized, and the studies they draw on—from Uher's treatment stratification to Bringmann's idiographic dynamics—show that those signals get louder when we listen symptom by symptom, over time, in context.
That's not just more precise science. It's a better map for helping the person in front of you.
Related lectures
- Insomnia and the risk of depression: a meta-analysis of prospective cohort studies
- Mirror-Induced Behavior in the Magpie (Pica pica): Evidence of Self-Recognition
- The cross-national epidemiology of social anxiety disorder: Data from the World Mental Health Survey Initiative
- The Natural Statistics of Audiovisual Speech
- The Small World of Psychopathology
- Health-related quality of life in parents of school-age children with Asperger syndrome or high-functioning autism