Start here: what major depressive disorder is
Major depressive disorder is diagnosed when persistently low mood or a loss of interest and pleasure — anhedonia — lasts for weeks, is accompanied by changes in sleep, appetite, energy, concentration and self-worth, and interferes with someone's ability to work, study or maintain relationships. It is not ordinary sadness, and it is not a failure of effort or character. It is common, it is disabling, and it is treatable, though less reliably than most people are led to believe.
The scale is not in dispute. As the authors of the dose-response meta-analysis at the centre of this review put it in their opening line, "depression is the single largest contributor to non-fatal health loss worldwide" [1], a position consistent with the Global Burden of Disease estimates across 204 countries [2] [3].
It is heterogeneous, and there is no single cause. Two people can both meet the diagnostic criteria while sharing only a minority of symptoms, because the criteria are polythetic — a threshold count from a list. That heterogeneity is not merely descriptive. Reviewing the biological dysregulations associated with depression, one analysis found that "the heterogeneity of the depression concept seems to play a differentiating role": metabolic syndrome and inflammatory up-regulation appear more specific to the atypical subtype, while hypercortisolaemia appears more specific to melancholic depression [4]. Different symptom profiles carry different biology, which is one reason averages across the diagnosis can be misleading.
A word about the "chemical imbalance" story. The idea that depression is caused by a shortage of serotonin has been repeated so often in public communication that many people believe it is settled. It is not. A systematic umbrella review synthesising the principal areas of relevant research — serotonin and 5-HIAA concentrations in body fluids, 5-HT₁ₐ receptor binding, serotonin transporter levels by imaging and post-mortem, tryptophan depletion studies, and SERT gene association and gene-environment interaction studies — found no consistent evidence that depression is associated with lowered serotonin concentration or activity. Two meta-analyses of the metabolite 5-HIAA showed no association with depression; a meta-analysis of cohort studies of plasma serotonin showed no relationship, and found instead that lowered serotonin was associated with antidepressant use [5].
That review has been argued with, and it does not show that antidepressants do not work — a drug can help without the condition being a deficiency of what the drug acts on, exactly as paracetamol helps a headache that is not caused by low paracetamol. But it does mean the simple deficiency story should not be told as established fact, and this review does not tell it.
Why honest numbers matter here specifically. Depression research is unusually vulnerable to three problems at once: outcomes are self-reported rather than measured by any biological test, placebo responses are large, and publication has historically been selective. Those three together can inflate an apparent effect considerably. Everything below therefore states effect sizes with the caveats attached, and where the literature disagrees, this review says so instead of picking a side.
Three pillars follow — measurements, treatments, and progress — with a dose-response model between the first two that carries the single most practically useful finding in the field.
Pillar 1: measurements and diagnosis
There is no test
This is the first thing to say plainly, because it shapes everything else. There is no blood test, no scan and no genetic assay that diagnoses major depressive disorder. Diagnosis is clinical: a trained interviewer establishes which symptoms are present, for how long, and how much they interfere with functioning. Severity is then tracked with rating scales — questionnaires either filled in by the patient or scored by a clinician during an interview.
The common instruments are the PHQ-9, a nine-item self-report scale scored 0–27 that maps directly onto the diagnostic criteria and is widely used in primary care [6] [7]; the clinician-rated Hamilton Depression Rating Scale (HAM-D), which dominates the older trial literature; the MADRS, used in most modern drug trials [8]; and the self-reported QIDS-SR, used in the sequential-care trial discussed below [9].
What a scale score does and does not mean
A scale score is a summary of what someone reports about the past two weeks. It is genuinely useful — it makes change visible, it can be repeated cheaply, and it turns a vague clinical impression into something trackable. It is also not a measurement of an underlying quantity in the way a blood pressure is.
Two properties matter for reading any trial.
The minimal clinically important difference. For the PHQ-9, in a study of 434 patients that assessed responsiveness, test–retest reliability and the smallest change that means anything, the minimal clinically important difference for individual change was estimated at 5 points on the 0–27 scale [10]. That number is the yardstick against which trial results should be read. A statistically significant two-point difference between groups is not the same thing as a change a patient would notice.
What "response" and "remission" mean. These are conventions, not natural categories. Response almost always means a 50% or greater reduction in symptom-scale score [1]; remission means falling below a defined low threshold [11] [9]. A person who moves from severe to mild depression counts as a responder while still being unwell.
The placebo response, and why it makes efficacy hard to measure
In depression trials the placebo group improves substantially. Some of that is genuine drug-independent improvement — attention, expectation, the therapeutic relationship, regression to the mean — and some is the natural course of an episode that would have lifted anyway. Because the placebo group improves so much, the difference between drug and placebo is a modest slice of a large total improvement, and it is that slice, not the total, that the drug can claim.
This is why the outcome definition matters so much, and why the field has taken publication bias seriously: when only favourable trials are published, a modest true effect looks larger. The CBT literature below is the clearest worked example, because its authors quantified their own bias correction [12].
Centerpiece: a simple simulatable model of the antidepressant dose-response
If a first antidepressant dose does not work well enough, the intuitive next move is to raise it. The best available evidence says that intuition is usually wrong, and the shape of the curve says why.
A systematic review and dose-response meta-analysis pooled 77 double-blind randomised trials of fixed doses — 19,364 participants, mean age 42.5 — of five SSRIs plus venlafaxine and mirtazapine, converting all doses to fluoxetine equivalents. Efficacy was treatment response, defined as a 50% or greater reduction in depression severity, after a median of 8 weeks [1]. Its central finding, in the authors' own words:
For SSRIs, the dose–efficacy curve "showed a gradual increase up to doses between 20 mg and 40 mg fluoxetine equivalents, and a flat to decreasing trend through the higher licensed doses up to 80 mg fluoxetine equivalents", while dropouts due to adverse effects "increased steeply through the examined range" [1].
Three curves, one dose axis: benefit that saturates early, harm that does not, and an optimum where the gap between them is widest.
What the model explains. Four things.
First, why "the dose isn't high enough" is usually the wrong explanation for a poor response. Most of the achievable benefit arrives by the bottom of the licensed range. Beyond the plateau the efficacy curve is flat to slightly declining while adverse-effect dropouts keep rising, so raising the dose past it trades a real increase in side effects for no measurable gain [1]. The same paper found the pattern held for venlafaxine, whose efficacy increased up to around 75–150 mg and then only modestly, and for mirtazapine, whose efficacy increased to about 30 mg and then decreased.
Second, why "optimal" is not the same as "maximum tolerated". The acceptability curve — dropouts for any reason, combining benefit and burden — was optimal in the lower licensed range, 20–40 mg fluoxetine equivalents [1]. The best dose is the one where the two curves are furthest apart, and that is near the bottom of the licensed range, not the top.
Third, why sequential care has sharply diminishing returns. The right panel is the other half of the picture and needs no model. Remission fell from 36.8% at the first treatment step to 13.0% at the fourth, and patients who required more steps relapsed more often afterwards [9]. The cumulative 67% is genuinely encouraging — most people who stay in care eventually remit — but the per-step yield collapses, and each step costs months. Later steps included switching drug and augmentation strategies, which were studied separately in the same programme [13].
Fourth, why this is the most actionable finding in the review. Unlike most of what follows, it is not contested, it is directly usable, and it points toward less intervention rather than more.
What the model deliberately does not do. The left panel's curves are shapes, not fitted values — the source reports splines and the review says so. It describes averages across trials at 8 weeks and cannot tell an individual whether their dose is right; some people genuinely do respond to higher doses, and the flat average curve is compatible with that. It says nothing about duration, which is a separate question from dose. And the right panel is a single large trial with a naturalistic design, whose generalisability and analysis have both been discussed extensively in the years since.
Pillar 2: treatments, with the numbers attached
Antidepressants: real but modest, and the certainty is not high
The largest synthesis is a network meta-analysis of 21 antidepressants, pooling 522 double-blind randomised trials and 116,477 participants, including unpublished data from regulatory agencies and company registries. Its headline: all 21 were more effective than placebo, with odds ratios for response ranging from 2.13 for amitriptyline down to 1.37 for reboxetine [14].
Three qualifications from the same paper belong in the same breath. Differences between active drugs were small when placebo-controlled trials were included, with wide credible intervals on most comparisons. For acceptability, only agomelatine and fluoxetine had fewer dropouts than placebo. And on quality: 9% of the 522 trials were rated at high risk of bias, 73% moderate, and the certainty of the evidence was moderate to very low [14].
So: antidepressants work, on average, better than placebo, by an amount that is real and clinically worthwhile for many people but is not large, and the evidence base carries acknowledged weaknesses. Both halves of that sentence are supported by the same paper, and quoting either half alone misrepresents it.
Adverse effects are common and matter for adherence — the dose-response curve above is largely a story about them [1] — and antidepressant use in older people carries its own risk profile [15]. Stopping is also not always straightforward: discontinuation symptoms are real, and the acceptability data in the network meta-analysis capture only dropout during acute trials, not the experience of coming off after months or years.
Psychotherapy: comparable, with a candid bias correction
A meta-analysis of 115 studies of cognitive-behavioural therapy for adult depression found a mean effect size across 94 comparisons with control groups of Hedges g = 0.71 — corresponding to a number needed to treat of 2.6. The authors then did something unusually forthright: they reported that this was probably an overestimate. After adjustment for publication bias the effect fell to g = 0.53, and higher-quality studies gave g = 0.53 against g = 0.90 for lower-quality ones [12].
Their conclusions are worth quoting for their restraint: there is no doubt CBT is effective, "although the effects may have been overestimated until now"; they found no indication that CBT was more or less effective than other psychotherapies or pharmacotherapy; and combined treatment was more effective than pharmacotherapy alone (g = 0.49) [12].
That last set of findings is the practically important one. For many people the choice between medication and psychotherapy can reasonably be made on preference, access and side-effect tolerance rather than on an expectation that one is clearly stronger. Other approaches have their own evidence — behavioural activation was compared against CBT on cost and outcome [16], internet-delivered CBT has been evaluated in its own right [17] [18], and CBT for comorbid insomnia improved depression outcomes [19]. Exercise has been meta-analysed with explicit publication-bias adjustment [20], and physical-activity interventions reviewed at the level of overviews [21].
Treatment-resistant depression
When two or more adequate trials of antidepressants fail, the condition is conventionally called treatment-resistant, and its definition, prevalence, detection and management have been reviewed as a distinct problem [22] [23]. This is the population in which the more intensive options are used.
Electroconvulsive therapy remains the most effective acute treatment for severe depression and has been assessed in systematic review and meta-analysis [24], with stimulus intensity and electrode placement both affecting the balance of efficacy against cognitive side effects [25]; relapse after a successful course is high enough that continuation pharmacotherapy afterwards has been trialled specifically [26]. Repetitive transcranial magnetic stimulation is less effective but far better tolerated, and has been studied since the mid-1990s [27].
Pillar 3: progress
Rapid-acting agents
The observation that changed the field's timescale was that a single intravenous dose of ketamine, an NMDA-receptor antagonist, improved depression within 110 minutes in patients with treatment-resistant major depression, with a very large drug–placebo effect size at 24 hours (d = 1.46) falling to moderate-to-large at one week (d = 0.68); 71% met response criteria the day after infusion and 35% maintained response for at least a week [28]. That trial had 18 subjects — the finding is important because conventional antidepressants take weeks, not because the trial was large.
Esketamine nasal spray took the mechanism into phase 3. In treatment-resistant depression, esketamine plus a newly initiated oral antidepressant beat an active comparator plus placebo spray, with a difference in MADRS change at day 28 of −4.0 points (95% CI −7.31 to −0.64) [8], with supporting fixed-dose [29] and adjunctive [30] [31] trials and a relapse-prevention study [32]. Note the size of that difference and its confidence interval: real, and modest. Dissociation, nausea, vertigo, dysgeusia and dizziness were all more frequent with esketamine, appearing shortly after dosing and generally resolving within 1.5 hours [8]. International expert opinion has synthesised the ketamine and esketamine evidence together [33], and mechanistic work continues on what beyond the NMDA receptor is involved [34] [35].
Psychedelics: promising, and not yet what the coverage suggests
Psilocybin has produced striking open-label results in treatment-resistant depression [36] and positive randomised results against waiting-list control in major depressive disorder [37] [38].
The most informative trial is the one that compared it against a real alternative. In a phase 2 double-blind randomised trial, 59 patients received either two 25 mg doses of psilocybin plus daily placebo, or two 1 mg doses plus daily escitalopram; everyone received psychological support. On the primary outcome — change in QIDS-SR-16 at week 6 — the psilocybin group improved by 8.0 points and the escitalopram group by 6.0, a between-group difference of 2.0 points that was not statistically significant (95% CI −5.0 to 0.9, P = 0.17) [11].
That is a genuinely interesting result and it is not a demonstration of superiority. Response rates favoured psilocybin numerically (70% vs 48%) with a confidence interval spanning no difference. Ayahuasca has also been tested in a randomised placebo-controlled trial in treatment-resistant depression [39], and long-term follow-up of psilocybin-assisted therapy exists in other populations [40]. The honest position is that this is an active and promising research area with small trials, difficult blinding — participants can usually tell which arm they are in — and no established place in routine care.
Measurement-based care, and predicting who responds
Two quieter developments may matter more than any new molecule. Measurement-based care — actually administering a scale at each visit and adjusting treatment on the result, rather than relying on clinical impression — is what makes the dose and step logic above usable, and it is why the responsiveness properties of the PHQ-9 were characterised in the first place [10] [6].
Predicting who responds remains largely unsolved. Work on serotonin-transporter gene methylation as a predictor of antidepressant response [41] and on psychiatric pharmacogenomics [42] illustrates the effort; nothing from it is yet reliable enough to choose a first-line treatment for an individual. Given the heterogeneity described at the top, this is the field's central open problem: the average effects reported throughout this review conceal people who benefit greatly and people who do not benefit at all, and at present there is no way to tell them apart in advance.
A closing note on what this review is not. Nothing here is advice, and none of the numbers describe any individual. Depression is treatable, most people who stay in care do improve, and the modesty of the average effect sizes is an argument for measurement, patience and choice among options — not for doing nothing. If you are struggling, the useful step is a conversation with a clinician, not a calculation.
Dig deeper in lmmol
Depression intersects with several conditions covered elsewhere in this series:
- Stroke — depression after stroke is common enough to be a named clinical problem, and the two reviews share the difficulty of measuring outcomes that matter to patients rather than only to scanners.
- Alzheimer's disease — depressive symptoms can be a risk factor for, and an early feature of, neurodegenerative disease, which makes the two diagnoses genuinely hard to separate in older adults.
- Heart failure and chronic kidney disease — chronic illness and depression are bidirectionally associated, and both reviews describe populations in which depression is common and frequently untreated.
- Obesity — another condition where public discourse routinely outruns the evidence, and where honest effect sizes matter for the same reasons.
- The health reviews index collects the rest of the series.
Then move down into lmmol's graph, to the proteins the drug classes act on — with the caveat from the opening section firmly attached, that acting on a target is not evidence that the target was the cause:
- Sodium-dependent serotonin transporter — SERT, the SSRI target, and one of the systems examined and found wanting as an explanation of depression itself [5] [41].
- Sodium-dependent noradrenaline transporter — the second target of the SNRIs.
- 5-HT₁ₐ and 5-HT₂ₐ receptors — the first examined in the serotonin umbrella review [5], the second the principal target of psilocybin [11].
- NMDA receptor subunit GluN1 and GluN2B — where ketamine and esketamine act [28] [34].
- Tryptophan 5-hydroxylase 2 — the rate-limiting enzyme of brain serotonin synthesis, and the pathway probed by the tryptophan-depletion studies in that umbrella review [5].
- Monoamine oxidase A — the target of the oldest antidepressant class [43].
- For entities without a linked static page here, use the graph index, all proteins, or all diseases rather than guessing an entity URL.