This comment is far more heat than light. Such strong claims need to be accompanied by some serious evidence. Consider this one week ban a warning should you wish to comment in the future.
What about pramipexole, which has demonstrated an effect size of 0.87 in a recent trial rather than 0.3. Maybe we just need better antidepressants ? 10.1016/S2215-0366(25)00194-4 I aggree of course with everything you say, that we need to leverage non pharmacological axis more
If you look at the supplement, the study had almost complete functional unblinding. It’s hard to make heads or tails of a study when this happens, when we already know non-pharmacological effects are very important. Substack won’t let me paste the image.
The statin comparison in your title may be doing more work than it first appears, and it points at something underneath the additivity debate in these comments. The question isn't only whether you can subtract the placebo arm. It's what part of the intervention actually travels.
A statin can leave the trial protocol and stay mostly itself. The molecule does the work, and the molecule is portable. What your data suggest about antidepressants is that the portable part, the prescription, is the smaller part, and the less portable parts, the contact schedule, the expectancy, the repeated chances to revise the story of what's happening, are where most of the effect was produced. Those aren't the setting the drug effect sits inside. In the trial they were partly constitutive of it.
So when practice exports the pill and drops the visit schedule, it isn't a lighter version of the trial. It's a different intervention still borrowing the trial's outcome language. "Six out of ten improve" reads as a property of the drug, but in the study it was a property of the drug inside an arrangement of care that primary care doesn't reproduce. This is a distinct claim from the ones being argued above about whether the effect is additive or heterogeneous, and I think it survives either answer: however the arithmetic resolves, the thing that got approved and the thing that gets prescribed are not the same object. One was studied inside a protocol. The other has been amputated from it.
Funny that you mention statins as a comparison, given the ways they so heavily advertise relative risk reduction, without even mentioning absolute risk reduction.
In that sense, psychiatry is not too different from the rest of medicine.
Great piece! Why it doesn’t explain everything. Catherine Harmer has a nice hypothesis/theory which ties in neurobiological/neuro cognitive + the extra factors in to explain SSRI effects https://pubmed.ncbi.nlm.nih.gov/31938879/
Non-pharmacological factors are clearly important in MDD treatment response. However, I don’t think your 80/20 interpretation follows from the small magnitude of trial-level mean differences. In particular I think that you’re making an implicit assumption of homogeneity of response.
The clearest counterargument actually comes from the Stone et al (2022) IPD meta-analysis that you cite and include a figure from at the top of your post. The distributions of HAMD changes in the figure are clearly not unimodal. Indeed, it is visually apparent that large improvements are more probable in the drug than the placebo group. Stone et al and other researchers have modelled this as a tri- (or elsewhere bi-) modal distribution.
Quoting Stone et al directly:
"Those treated with drug were more likely to show a Large response (24.5% v 9.6% with placebo), however, and less likely to have a Minimal response (12.2% v 21.5%). Thus the observed advantage of antidepressants over placebo is best understood as affecting a minority of patients as either an increase in the likelihood of a Large response or a decrease in the likelihood of a Minimal response."
In short, and to quote a similar study by Thase et al (2011, BrJPsych):“small mean differences obscure large and clinically meaningful responses for a subgroup of people with depression.” This is why, as a clinician, I am generally more interested in the NNT for response and remission.
Thanks for this comment. Nils and I have argued about this aspect of the paper. A few places where I would respectfully disagree.
First, the response rates in the antidepressant arm reflects primarily non-pharmacological factors. We cannot pick the tail end of the distribution (and, say, call it a “large response”) and argue that those patients’ responses were drive by pharmacological factors. There is also a distribution of responders with placebo, with a tail end of super responders to placebo, as is clearest in the esketamine figure I presented. There is just no greater variation in antidepressant response than with placebo: https://jamanetwork.com/journals/jamapsychiatry/fullarticle/2776611.
Across the entire sample, the average difference in improvement was 8 vs 10 HAMD points. Any medication patients who had a greater improvement relative to placebo than this, means that other medication patients had less improvement than this. Does this risk the ecological fallacy? Sure, but we have no alternative when discussing individuals. There have been endless studies attempting to differentially predict responders to antidepressant vs placebo…and none to my knowledge are robust. That’s where I disagree with the Thase quote. Just because there is a distribution of response, with slightly more patients doing slightly better at every threshold of improvement, does not mean that we can intelligently determine who will be a large responder or whether the risk/benefit justifies this. There’s lots of evidence against this idea.
Second, response arbitrarily dichotomize continuous data to inflate small differences. This has been widely shown. Stone does their own version of this by grouping people in to arbitrary categories of “large” and “minimal” responders. But these are not natural entities. The modeling of distributions into tri or bimodal distributions are disputed https://www.bmj.com/content/378/bmj-2021-067606/rapid-responses, specific models do not typically replicate (e.g. with STAR*D https://www.sciencedirect.com/science/article/pii/S0895435625002768), and apply largely arbitrary categories. And again, I emphasize that the bulk of improvement can only be attributed to (various) non-pharmacological factors in who they classify as “large responders.” They are not predicting individual response or identifying who is a med responder.
Maybe you find their statistical modeling compelling. I’m very suspicious. There’s just not a good track record of these models being robust. Just my two cents.
In fairness, Thase and others wrote a good paper on this subject that reflects the points you make called Practising evidence-based medicine in an era of high placeboresponse: number needed to treat reconsidered
I agree that the NNT can inflate small differences, but any summary statistic loses information compared to the full distribution. In the case of the mean difference, again, paraphrasing Thase, it can obscure large, clinically meaningful differences in a subgroup of participants. Inspecting the Stone et al response distribution, you could introduce a dichotomy at a number of points (HAMD reduction >10, or >15, or >20, etc) and a clear NNT advantage for pharmacotherapy would still be shown. In this instance, it's plainly not an artefact of an arbitrary threshold or the choice to use Gaussian modelling (even though examples can be *constructed* that fall prey to such pathologies). Large responses are clearly more probable with antidepressant treatment across a range of thresholds in the Stone IPD dataset. And this matters clinically.
I agree that we lack tools to robustly predict antidepressant response, but that is beside the point here. Nonetheless, I note that many factors have been studied that *may* account for heterogeneity in treatment response, including inflammation (indexed by CRP), BMI, trait neuroticism, early life adversity, and so on.
The Thase quote isn’t claiming that we can identify in advance who will have a large response. It simply claims that such a subgroup exists. For now, evidence-guided trial-and-error is unavoidable. In practice, I pursue sequential, time-limited trials of medications with distinct mechanisms (alongside lifestyle measures and psychotherapy, of course). I conceptualise this approach as an evidence-guided search process, and it is supported by the various RCTs on so-called algorithmic prescribing and measurement-based care.
And yes, in my opinion it is an ecological fallacy to claim that an 8 vs 10 point difference implies that 80% of improvement is due to non-pharmacological factors in all individuals. It’s reasonable to hypothesise this, and to explore possible mechanisms, but it isn’t a conclusion that is supported by the data. Also note that a difference in variance between groups is only one signature of response heterogeneity, and the absence of such a difference cannot exclude heterogeneity.
Within reason, the distribution of outcomes for drug and placebo are extremely similar, just shifted ~2 points. The idea of a subgroup of large responders is just not there. There's a relatively smooth distribution at each improvement, just slightly larger for drug vs placebo.
I tried to convert the Stone bar chart to a table (so it may be imperfect). Of patients having at least 20 HAMD points improvement, this was 15.3% with med and 11.4% with placebo. That's a NNT of 25. It ranges around 15-25 for various cutpoints, as relatively few patients had large improvements with med or placebo. Michael Hengartner has written extensively about different ways to model the overlap in these trials. And all of this ignores reasonable critiques of these trials, like effective unblinding. Recent papers assessing this suggest high rates of effective unblinding, e.g. the PANDA trial or escitalopram vs tDCS.
It just surprises me that there's so much discomfort with non-pharmacological factors being important in depression (note, I'm not arguing about OCD or psychosis), when there's lots of evidence in favor of this, and very little falsifiable evidence to the contrary. Many studies seem hard to explain through this model. The olanzapine-fluoxetine combination trials (https://thinkingpsychiatry.substack.com/p/the-most-important-study-in-depression), some of the few multiple active-arm trials with a placebo-equivalent control in psychiatry. STAR*D level 2 switch showed equivalent response after citalopram failure to sertraline, bupropion, or venlafaxine, meaning that responders to each med were perfectly balanced by chance. Double blind RCTs of switching, dose increases, blinded combination trials, iSPOT-D. Equivalent response across basically every antidepressant med/class. I've written about the data on MBC, as well, through this lens. Lots of examples.
My main frustration with the Thase line of reasoning is that it seems to justify any treatment, with any med, because there's always the possibility that you could be a responder--despite virtually no meaningful prospective evidence after decades of study (I'm referring to models of response based on drug mechanism). It exists with antipsychotics routinely when we have evidence of 95% effective doses which, by definition, mean that only 5% of people can get benefit beyond that dose. Ironically, even STAR*D explicitly argued against this model in the level 3 switch and augmentation papers: "This study highlights the need for strategies beyond following a simple sequence of treatments."
Thanks for the reply, Kevin. We'll have to agree to disagree on this one. I don't judge the two distributions in the Stone IPD paper to be the same but shifted by 2 points. Both distributions are bimodal, but to me eye (and this is supported by the Gaussian modelling) the medication distribution has a larger left hump while the placebo distribution has a larger right hump. Again, this asymmetry, which I still contend is clinically meaningful, is obscured if we simply compare the means.
Perhaps a speculative view but I think much of depression relates to rigid internal self-world models (with likely multiple different aetiologies leading to same) and that SSRIs may be working by reducing the rigidity of these models allowing flexible updating, similar to the REBUS model - just chronically rather than acutely. If SSRI are facilitating belief updating on a biological level, then the expectancy effect is literally part of treatment itself - functioning to update models at a cognitive level. Then take someone in depressive withdrawal and bring them to a bunch of appointments etc and you are furthermore increasing the possibility of model updating through the relational/behavioural aspect. I'm not sure any of these effects are cleanly separable from each other if they are all modulating rigid predictive models from different angles (either from inside out or outside in). Whilst hypothetical, there does seem to be a signal here (https://www.nature.com/articles/s44220-026-00621-9).
Also, the actual aetiology of depression could be (and likely is) changing which could well reflect different placebo/non-pharmacological response rates than previous. Perhaps we are dealing with much less primary biological depression and more psycho-socially mediated, in which case it seems response would be better to non-pharmacological aspects than previous.
This is a compelling model. I’m familiar with other biological models of SSRIs around either synaptic plasticity or hippocampal neurogenesis as creating environmental synergies like you describe.
My hesitation with this is that none of this appears unique to SSRIs. Why would this work for mirtazapine, bupropion, quetiapine 50mg, and so forth where these shouldn’t be the mechanism of action? Why would there be equivalent outcomes in STAR*D level 2 switch, in the olanzapine-fluoxetine combination trials, or basically equivalent outcomes across every class in antidepressant meta-analyses?
I also become more hesitant when each new branded medication revises our biological explanatory model, scopolamine infusions being a key example.
My money is on expectancy (which psychiatry helped create) and baseline characteristics/natural history
As I say I'm being somewhat speculative and continuing be so, I think it's possible different medications are driving this 'model-updating' through different mechanisms. For example, buproprion is activating and could almost be seen as a pharmaceutical aid to behavioural activation and so catalyse model updating through engagement with the world rather than withdrawal.
There may be other pharmacological routes to restoring flexibility, not specific to SSRI - either through altering salience, restoring sleep/reducing the re-inforcing of beliefs through rumination/distress etc.
Scolpamine, Ketamine, psychedelics, ECT - all through different mechanisms seem to be involved in network level disruption and neuroplasticity. The common potential mechanism there to me seems to be disruption of stuck 'modelling' and facilitation of updating of those models.
Of course that doesn't discount the role of expectation - what is expectation other than a belief?
I think it's very reasonable for different medications to act within different ways--to an extent. The problem I have is that 1) there needs to be some plausible limit to the theoretical models we have and 2) they need to be falsifiable.
I mean, the list of treatments with positive trials is just comically long at this point, botuilinum toxin, nitrous oxide, inlahed sevoflurane, propofol. Scopolamine IV is an interesting example because it had large effect sizes versus placebo, but failed vs an active control.
A few examples come to mind involving bupropion. Level 2 of STAR*D involved a switch from people who failed citalopram (true trials) to sertraline, venlafaxine, or bupropion. There were equivalent outcomes, which suggests either a perfect balance of these various mechanisms by chance, or another explanatory mechanism. Bupropion has equivalent outcomes on anxiety in depression as SSRIs, even though this contradicts our model of mechanism. I've written about the olanzapine-fluoxetine combination trials which are the clearest example of this: https://thinkingpsychiatry.substack.com/p/the-most-important-study-in-depression. Curious your thoughts on those large studies.
It feels like the parsimonious explanation is some non-pharmacological mechanism--not saying we know what this involves or that this is simple.
Anecdotally, Irving Kirsch's interests in antidepressants came from this model: if hopelessness is a core symptom of depression, and a medication embodies some aspect of hope and expectancy, does the prescription itself therefore treat a core depression symptom independent of mechanism.
I agree there needs to be limits, but if a predictive coding/rigid belief account of depression is right, it’s going to be more complex causatively than one clear pathway given the number of different processes involved, any of which could lead to downstream model rigidity as the common factor (including psychosocial factors). If we know the specific causative mechanism then we can obviously target that, but even without there are likely many ways to influence modelling processes and reduce rigidity.
I think your ideas about the expectancy effect fit very well with this. What is hopelessness other than the stuck belief that 'things are never going to get better'? External intervention becomes necessary because those rigid beliefs are likely shaping perception/action such that subtle evidence to the contrary may be ignored/inaccessible, re-inforcing the 'stuckness'. A psychiatrist saying here, take this pill it'll make you feel better is as you say - providing hope and expectancy, or challenging that stuck belief directly, regardless of what we are prescribing. Essentially we are introducing a prediction error that says: “hold up - maybe there is hope?”. Kind of like how in trauma therapy there is relational safety challenging threat based models, but here relational hope.
That seems in keeping with the power of non-pharmacological effects you mention in the linked article (the venlafaxine example is really quite telling). I don't think that entirely disregards the possibility of these medications having (perhaps more subtle) effects on belief rigidity though, through independent biological mechanisms.
How to falsify this, I'm not sure and I am certainly no academic, but could you consider whether treatment modulates measures of belief rigidity (i.e. cognitive flexibility or computational modelling studies) and is this modulation associated with reduced depression scores and to what extent does this differ between active/placebo arms? On Bupropion specifically I would be intrigued to see how effective it is combined with behavioral activation vs as a sole pharmacological agent without any behavioural modification. Also the evidence that for severe depression SSRI + psychotherapy is supportive.
If the expectancy effect is such a big player - I wonder how damaging it is to label something ‘treatment resistant’... rather than giving hope, reinforcing the hopelessness…
Really appreciate this framing; I talk about this all the time with patients (mostly around TMS). I think we can be honest about it without either underselling our treatments (by saying there’s no difference from PBO, which isn’t true) nor over promising (by emphasizing the biochemical effects). The act of coming in for help, talking to someone of authority, getting a coherent story about why you’re experiencing this and how we’re going to try to help, and coming back to talk about it some more - all very powerful stuff that adds to the (admittedly small) direct physiological effects we know are happening from the PBO/sham difference
This is put really well, and is the exact framing that I was trying to get at.
The place I would gently disagree is the extent to which how we (collectively) portray treatments, their mechanism, and the cause of depression generates expectancy that require our treatments to then tap into. As another comment mentioned, expectancy (a traditional placebo effect) is probably doing a lot of the work, along with baseline characteristics/natural history. But we as a field shape those expectations. Fifty years ago, there weren't popular TV shows with musicals about how taking an antidepressant is no big deal because everyone takes one.
I don't share your skepticism. Having been on antidepressants myself and having multiple family member on them, I am extremely amenable to the idea that most of the effect is non-pharmacological. There is significant relief to the user just from knowing. "Someone is looking out for me. This thing will help me." And for an illness that primarily works by sapping hope, that effect is extremely useful.
At the same time, my observation from the exterior is that the recipients seem not to have improved in physically measurable ways (energy levels, etc) even when they report psychological benefits.
Of course, this is all my personal experience and instincts, just like your skepticism comes from yours.
In the end, depression is a social/psychological problem. It makes sense the the best cures would be too.
Excellent work. One friendly push on the arithmetic, because I think it may strengthen your conclusion rather than weaken it.
The 80/20 framing treats the drug effect and placebo response as separable line items: total improvement minus the placebo arm equals the “real” drug effect. But that implies additivity. Benedetti’s placebo work suggests something messier: expectancy and ritual recruit measurable endogenous opioid, dopaminergic, serotonergic, and stress/modulatory systems. As you suggest, context is a biological input rather than noise.
The rTMS placebo literature points the same way (Razza et al. found a large sham response in depression trials that was associated with improvement in the active-treatment group, suggesting placebo response may be part of therapeutic response rather than merely a confound). https://pubmed.ncbi.nlm.nih.gov/29111404/
So the measured 2-point drug-placebo difference may not be a fixed pharmacologic payload sitting on top of an 8-point contextual effect as much as the drug’s marginal effect within an already powerful clinical context.
That matters for the lower-dose suggestion. Using lower doses to leverage placebo while reducing risk assumes dose and context are independent levers. If they interact, the trade may not hold.
I really liked your suggestion that clinicians can likely get more treatment benefit clinically by spending less time optimizing drug chemistry and more time on the "extra" context factors like trust, warmth, attention, and frequency of visits.
I don’t think the recommendation follows. The article shows placebo groups improve a lot, but that doesn’t show the improvement comes from warmth, follow-ups, hope-building, or behavioral activation.
Can you really derive the magnitude of the medication-only effect by subtracting the effect size of the placebo arm? Seems possible the effects of psychological expectation, social treatment and role, and pharmacology are not independently additive. You'd have to have something like a trial of people running placebo-controlled experiments at home to figure this out, but it seems very possible to me that the pharmacology alone could knock out some of the same 'depression points' that are covered by being attended to weekly by a kindly psychiatrist etc. I forget what this is called statistically but it's like the opposite of gestalt, the whole is less than the sum of its parts.
The study design you mention is the one I referenced by Leuchter, as being the best effort at this, but it unfortunately had big baseline differences in the exact thing they were trying to study.
The placebo response is agnostic about the source of improvement with placebo. We just don't have a great sense of which factors matter most, how they overlap, whether they are additive. Personally, I think baseline characteristics/natural history drive most outcomes, with expectancy and contact next.
What you seem to be getting at is the idea that prescribing an active medication, in clinical practice, would tap into other non-pharmacological factors like expectancy. Doing nothing, which is what placebo represents in the rest of medicine, wouldn't. And I totally agree with this. A great paper by Roose, Rutherford, and Thase makes a similar point to argue that the only way to generate real-world non-pharmacological factors is to prescribe a medication. They have presented some data that more frequent appointments has larger effects in the placebo than medication arm.
But I worry that your perspective makes the assumption that a prescription captures most of these non-pharmacological effects. I just don't think this assumption is fair to make or supported by evidence. I'm happy to be shown contrary evidence. This should be easy to prove with clinical trials of reduced contact and infrequent visits, and yet no one runs these trials.
My mental model isn't that the known factors might reduce to each other or be straightforwardly redundant, but rather that they might exert effects on the same nodes in a higher-order network that is currently undescribed. To make a toy example, if the factors delineated here (let's call them pharmacology, psychology, and social effects, or A, B, and C) all affect some three other unknown variables that are closer to the bone (call them 1, 2, and 3) it might be the case that A affects 1 and 2, and B affects 2 and 3, and C affects all three, but if 1 has already been pushed by A, you don't get further effect on 1 by also pushing C. So if you have a study arm that estimates the effects of B and C (like a placebo arm) and an arm that estimates the effects of A, B, and C (experimental arm) it wouldnt follow that the effect of A is approximately what's observed in the experimental arm - placebo arm.
Thank you for articulating this, as this was my initial reaction too. I won't deny the power of non-pharmacological/contextual effects on placebo-arm trial participants (speaking from personal experience participating in an ineffectively blinded study), but it strikes me as unlikely that it would be this simple.
I'm also skeptical that every single real-life patient who experiences meaningful improvement on an antidepressant derives the vast majority of that benefit from placebo response, especially if the average patient care experience is nothing more than a quarterly 10-minute check-in. Maybe the explanation is that responders receive an above-average level of care, but I don't know if that's been studied. Intuition of course doesn't always bear out in reality, but just offering my $0.02.
All fair points and my goal wasn't to suggest any of this is simple--I think it's really complicated, especially in practice.
Nils and I have argued about your point. This is where I feel like clinical trials offer the relevant counterfactual. There are just countless clinical trials, like the esketamine one I cite, that show massive improvement in the placebo arm.
The reality is that there is a range of response to both antidepressant and placebo, some people get a little better, some a lot better. But that's seen in both arms, like the figure from the Stone paper I showed.
The parsimonious explanation is that non-pharmacological seem to matter most in common disorders--at least in the clinical trials that, ironically, lead to FDA med approval.
This comment is far more heat than light. Such strong claims need to be accompanied by some serious evidence. Consider this one week ban a warning should you wish to comment in the future.
See depressed patients weekly, its that simple
What about pramipexole, which has demonstrated an effect size of 0.87 in a recent trial rather than 0.3. Maybe we just need better antidepressants ? 10.1016/S2215-0366(25)00194-4 I aggree of course with everything you say, that we need to leverage non pharmacological axis more
If you look at the supplement, the study had almost complete functional unblinding. It’s hard to make heads or tails of a study when this happens, when we already know non-pharmacological effects are very important. Substack won’t let me paste the image.
The statin comparison in your title may be doing more work than it first appears, and it points at something underneath the additivity debate in these comments. The question isn't only whether you can subtract the placebo arm. It's what part of the intervention actually travels.
A statin can leave the trial protocol and stay mostly itself. The molecule does the work, and the molecule is portable. What your data suggest about antidepressants is that the portable part, the prescription, is the smaller part, and the less portable parts, the contact schedule, the expectancy, the repeated chances to revise the story of what's happening, are where most of the effect was produced. Those aren't the setting the drug effect sits inside. In the trial they were partly constitutive of it.
So when practice exports the pill and drops the visit schedule, it isn't a lighter version of the trial. It's a different intervention still borrowing the trial's outcome language. "Six out of ten improve" reads as a property of the drug, but in the study it was a property of the drug inside an arrangement of care that primary care doesn't reproduce. This is a distinct claim from the ones being argued above about whether the effect is additive or heterogeneous, and I think it survives either answer: however the arithmetic resolves, the thing that got approved and the thing that gets prescribed are not the same object. One was studied inside a protocol. The other has been amputated from it.
Excellent work - thank you for this.
Funny that you mention statins as a comparison, given the ways they so heavily advertise relative risk reduction, without even mentioning absolute risk reduction.
In that sense, psychiatry is not too different from the rest of medicine.
Great piece! Why it doesn’t explain everything. Catherine Harmer has a nice hypothesis/theory which ties in neurobiological/neuro cognitive + the extra factors in to explain SSRI effects https://pubmed.ncbi.nlm.nih.gov/31938879/
Non-pharmacological factors are clearly important in MDD treatment response. However, I don’t think your 80/20 interpretation follows from the small magnitude of trial-level mean differences. In particular I think that you’re making an implicit assumption of homogeneity of response.
The clearest counterargument actually comes from the Stone et al (2022) IPD meta-analysis that you cite and include a figure from at the top of your post. The distributions of HAMD changes in the figure are clearly not unimodal. Indeed, it is visually apparent that large improvements are more probable in the drug than the placebo group. Stone et al and other researchers have modelled this as a tri- (or elsewhere bi-) modal distribution.
Quoting Stone et al directly:
"Those treated with drug were more likely to show a Large response (24.5% v 9.6% with placebo), however, and less likely to have a Minimal response (12.2% v 21.5%). Thus the observed advantage of antidepressants over placebo is best understood as affecting a minority of patients as either an increase in the likelihood of a Large response or a decrease in the likelihood of a Minimal response."
In short, and to quote a similar study by Thase et al (2011, BrJPsych):“small mean differences obscure large and clinically meaningful responses for a subgroup of people with depression.” This is why, as a clinician, I am generally more interested in the NNT for response and remission.
Thanks for this comment. Nils and I have argued about this aspect of the paper. A few places where I would respectfully disagree.
First, the response rates in the antidepressant arm reflects primarily non-pharmacological factors. We cannot pick the tail end of the distribution (and, say, call it a “large response”) and argue that those patients’ responses were drive by pharmacological factors. There is also a distribution of responders with placebo, with a tail end of super responders to placebo, as is clearest in the esketamine figure I presented. There is just no greater variation in antidepressant response than with placebo: https://jamanetwork.com/journals/jamapsychiatry/fullarticle/2776611.
Across the entire sample, the average difference in improvement was 8 vs 10 HAMD points. Any medication patients who had a greater improvement relative to placebo than this, means that other medication patients had less improvement than this. Does this risk the ecological fallacy? Sure, but we have no alternative when discussing individuals. There have been endless studies attempting to differentially predict responders to antidepressant vs placebo…and none to my knowledge are robust. That’s where I disagree with the Thase quote. Just because there is a distribution of response, with slightly more patients doing slightly better at every threshold of improvement, does not mean that we can intelligently determine who will be a large responder or whether the risk/benefit justifies this. There’s lots of evidence against this idea.
Second, response arbitrarily dichotomize continuous data to inflate small differences. This has been widely shown. Stone does their own version of this by grouping people in to arbitrary categories of “large” and “minimal” responders. But these are not natural entities. The modeling of distributions into tri or bimodal distributions are disputed https://www.bmj.com/content/378/bmj-2021-067606/rapid-responses, specific models do not typically replicate (e.g. with STAR*D https://www.sciencedirect.com/science/article/pii/S0895435625002768), and apply largely arbitrary categories. And again, I emphasize that the bulk of improvement can only be attributed to (various) non-pharmacological factors in who they classify as “large responders.” They are not predicting individual response or identifying who is a med responder.
Maybe you find their statistical modeling compelling. I’m very suspicious. There’s just not a good track record of these models being robust. Just my two cents.
In fairness, Thase and others wrote a good paper on this subject that reflects the points you make called Practising evidence-based medicine in an era of high placeboresponse: number needed to treat reconsidered
Thanks for the reply!
I agree that the NNT can inflate small differences, but any summary statistic loses information compared to the full distribution. In the case of the mean difference, again, paraphrasing Thase, it can obscure large, clinically meaningful differences in a subgroup of participants. Inspecting the Stone et al response distribution, you could introduce a dichotomy at a number of points (HAMD reduction >10, or >15, or >20, etc) and a clear NNT advantage for pharmacotherapy would still be shown. In this instance, it's plainly not an artefact of an arbitrary threshold or the choice to use Gaussian modelling (even though examples can be *constructed* that fall prey to such pathologies). Large responses are clearly more probable with antidepressant treatment across a range of thresholds in the Stone IPD dataset. And this matters clinically.
I agree that we lack tools to robustly predict antidepressant response, but that is beside the point here. Nonetheless, I note that many factors have been studied that *may* account for heterogeneity in treatment response, including inflammation (indexed by CRP), BMI, trait neuroticism, early life adversity, and so on.
The Thase quote isn’t claiming that we can identify in advance who will have a large response. It simply claims that such a subgroup exists. For now, evidence-guided trial-and-error is unavoidable. In practice, I pursue sequential, time-limited trials of medications with distinct mechanisms (alongside lifestyle measures and psychotherapy, of course). I conceptualise this approach as an evidence-guided search process, and it is supported by the various RCTs on so-called algorithmic prescribing and measurement-based care.
And yes, in my opinion it is an ecological fallacy to claim that an 8 vs 10 point difference implies that 80% of improvement is due to non-pharmacological factors in all individuals. It’s reasonable to hypothesise this, and to explore possible mechanisms, but it isn’t a conclusion that is supported by the data. Also note that a difference in variance between groups is only one signature of response heterogeneity, and the absence of such a difference cannot exclude heterogeneity.
A few thoughts.
Within reason, the distribution of outcomes for drug and placebo are extremely similar, just shifted ~2 points. The idea of a subgroup of large responders is just not there. There's a relatively smooth distribution at each improvement, just slightly larger for drug vs placebo.
I tried to convert the Stone bar chart to a table (so it may be imperfect). Of patients having at least 20 HAMD points improvement, this was 15.3% with med and 11.4% with placebo. That's a NNT of 25. It ranges around 15-25 for various cutpoints, as relatively few patients had large improvements with med or placebo. Michael Hengartner has written extensively about different ways to model the overlap in these trials. And all of this ignores reasonable critiques of these trials, like effective unblinding. Recent papers assessing this suggest high rates of effective unblinding, e.g. the PANDA trial or escitalopram vs tDCS.
It just surprises me that there's so much discomfort with non-pharmacological factors being important in depression (note, I'm not arguing about OCD or psychosis), when there's lots of evidence in favor of this, and very little falsifiable evidence to the contrary. Many studies seem hard to explain through this model. The olanzapine-fluoxetine combination trials (https://thinkingpsychiatry.substack.com/p/the-most-important-study-in-depression), some of the few multiple active-arm trials with a placebo-equivalent control in psychiatry. STAR*D level 2 switch showed equivalent response after citalopram failure to sertraline, bupropion, or venlafaxine, meaning that responders to each med were perfectly balanced by chance. Double blind RCTs of switching, dose increases, blinded combination trials, iSPOT-D. Equivalent response across basically every antidepressant med/class. I've written about the data on MBC, as well, through this lens. Lots of examples.
My main frustration with the Thase line of reasoning is that it seems to justify any treatment, with any med, because there's always the possibility that you could be a responder--despite virtually no meaningful prospective evidence after decades of study (I'm referring to models of response based on drug mechanism). It exists with antipsychotics routinely when we have evidence of 95% effective doses which, by definition, mean that only 5% of people can get benefit beyond that dose. Ironically, even STAR*D explicitly argued against this model in the level 3 switch and augmentation papers: "This study highlights the need for strategies beyond following a simple sequence of treatments."
Thanks for the reply, Kevin. We'll have to agree to disagree on this one. I don't judge the two distributions in the Stone IPD paper to be the same but shifted by 2 points. Both distributions are bimodal, but to me eye (and this is supported by the Gaussian modelling) the medication distribution has a larger left hump while the placebo distribution has a larger right hump. Again, this asymmetry, which I still contend is clinically meaningful, is obscured if we simply compare the means.
Perhaps a speculative view but I think much of depression relates to rigid internal self-world models (with likely multiple different aetiologies leading to same) and that SSRIs may be working by reducing the rigidity of these models allowing flexible updating, similar to the REBUS model - just chronically rather than acutely. If SSRI are facilitating belief updating on a biological level, then the expectancy effect is literally part of treatment itself - functioning to update models at a cognitive level. Then take someone in depressive withdrawal and bring them to a bunch of appointments etc and you are furthermore increasing the possibility of model updating through the relational/behavioural aspect. I'm not sure any of these effects are cleanly separable from each other if they are all modulating rigid predictive models from different angles (either from inside out or outside in). Whilst hypothetical, there does seem to be a signal here (https://www.nature.com/articles/s44220-026-00621-9).
Also, the actual aetiology of depression could be (and likely is) changing which could well reflect different placebo/non-pharmacological response rates than previous. Perhaps we are dealing with much less primary biological depression and more psycho-socially mediated, in which case it seems response would be better to non-pharmacological aspects than previous.
This is a compelling model. I’m familiar with other biological models of SSRIs around either synaptic plasticity or hippocampal neurogenesis as creating environmental synergies like you describe.
My hesitation with this is that none of this appears unique to SSRIs. Why would this work for mirtazapine, bupropion, quetiapine 50mg, and so forth where these shouldn’t be the mechanism of action? Why would there be equivalent outcomes in STAR*D level 2 switch, in the olanzapine-fluoxetine combination trials, or basically equivalent outcomes across every class in antidepressant meta-analyses?
I also become more hesitant when each new branded medication revises our biological explanatory model, scopolamine infusions being a key example.
My money is on expectancy (which psychiatry helped create) and baseline characteristics/natural history
As I say I'm being somewhat speculative and continuing be so, I think it's possible different medications are driving this 'model-updating' through different mechanisms. For example, buproprion is activating and could almost be seen as a pharmaceutical aid to behavioural activation and so catalyse model updating through engagement with the world rather than withdrawal.
There may be other pharmacological routes to restoring flexibility, not specific to SSRI - either through altering salience, restoring sleep/reducing the re-inforcing of beliefs through rumination/distress etc.
Scolpamine, Ketamine, psychedelics, ECT - all through different mechanisms seem to be involved in network level disruption and neuroplasticity. The common potential mechanism there to me seems to be disruption of stuck 'modelling' and facilitation of updating of those models.
Of course that doesn't discount the role of expectation - what is expectation other than a belief?
I think it's very reasonable for different medications to act within different ways--to an extent. The problem I have is that 1) there needs to be some plausible limit to the theoretical models we have and 2) they need to be falsifiable.
I mean, the list of treatments with positive trials is just comically long at this point, botuilinum toxin, nitrous oxide, inlahed sevoflurane, propofol. Scopolamine IV is an interesting example because it had large effect sizes versus placebo, but failed vs an active control.
A few examples come to mind involving bupropion. Level 2 of STAR*D involved a switch from people who failed citalopram (true trials) to sertraline, venlafaxine, or bupropion. There were equivalent outcomes, which suggests either a perfect balance of these various mechanisms by chance, or another explanatory mechanism. Bupropion has equivalent outcomes on anxiety in depression as SSRIs, even though this contradicts our model of mechanism. I've written about the olanzapine-fluoxetine combination trials which are the clearest example of this: https://thinkingpsychiatry.substack.com/p/the-most-important-study-in-depression. Curious your thoughts on those large studies.
It feels like the parsimonious explanation is some non-pharmacological mechanism--not saying we know what this involves or that this is simple.
Anecdotally, Irving Kirsch's interests in antidepressants came from this model: if hopelessness is a core symptom of depression, and a medication embodies some aspect of hope and expectancy, does the prescription itself therefore treat a core depression symptom independent of mechanism.
I agree there needs to be limits, but if a predictive coding/rigid belief account of depression is right, it’s going to be more complex causatively than one clear pathway given the number of different processes involved, any of which could lead to downstream model rigidity as the common factor (including psychosocial factors). If we know the specific causative mechanism then we can obviously target that, but even without there are likely many ways to influence modelling processes and reduce rigidity.
I think your ideas about the expectancy effect fit very well with this. What is hopelessness other than the stuck belief that 'things are never going to get better'? External intervention becomes necessary because those rigid beliefs are likely shaping perception/action such that subtle evidence to the contrary may be ignored/inaccessible, re-inforcing the 'stuckness'. A psychiatrist saying here, take this pill it'll make you feel better is as you say - providing hope and expectancy, or challenging that stuck belief directly, regardless of what we are prescribing. Essentially we are introducing a prediction error that says: “hold up - maybe there is hope?”. Kind of like how in trauma therapy there is relational safety challenging threat based models, but here relational hope.
That seems in keeping with the power of non-pharmacological effects you mention in the linked article (the venlafaxine example is really quite telling). I don't think that entirely disregards the possibility of these medications having (perhaps more subtle) effects on belief rigidity though, through independent biological mechanisms.
How to falsify this, I'm not sure and I am certainly no academic, but could you consider whether treatment modulates measures of belief rigidity (i.e. cognitive flexibility or computational modelling studies) and is this modulation associated with reduced depression scores and to what extent does this differ between active/placebo arms? On Bupropion specifically I would be intrigued to see how effective it is combined with behavioral activation vs as a sole pharmacological agent without any behavioural modification. Also the evidence that for severe depression SSRI + psychotherapy is supportive.
If the expectancy effect is such a big player - I wonder how damaging it is to label something ‘treatment resistant’... rather than giving hope, reinforcing the hopelessness…
This is a very thoughtful piece. I looked at changes in the nature of the clinical trials population over the last decades as one major factor in driving up placebo rates: https://drstevepsychiatry.substack.com/p/what-is-going-on-with-placebo-responses
Really appreciate this framing; I talk about this all the time with patients (mostly around TMS). I think we can be honest about it without either underselling our treatments (by saying there’s no difference from PBO, which isn’t true) nor over promising (by emphasizing the biochemical effects). The act of coming in for help, talking to someone of authority, getting a coherent story about why you’re experiencing this and how we’re going to try to help, and coming back to talk about it some more - all very powerful stuff that adds to the (admittedly small) direct physiological effects we know are happening from the PBO/sham difference
This is put really well, and is the exact framing that I was trying to get at.
The place I would gently disagree is the extent to which how we (collectively) portray treatments, their mechanism, and the cause of depression generates expectancy that require our treatments to then tap into. As another comment mentioned, expectancy (a traditional placebo effect) is probably doing a lot of the work, along with baseline characteristics/natural history. But we as a field shape those expectations. Fifty years ago, there weren't popular TV shows with musicals about how taking an antidepressant is no big deal because everyone takes one.
I don't share your skepticism. Having been on antidepressants myself and having multiple family member on them, I am extremely amenable to the idea that most of the effect is non-pharmacological. There is significant relief to the user just from knowing. "Someone is looking out for me. This thing will help me." And for an illness that primarily works by sapping hope, that effect is extremely useful.
At the same time, my observation from the exterior is that the recipients seem not to have improved in physically measurable ways (energy levels, etc) even when they report psychological benefits.
Of course, this is all my personal experience and instincts, just like your skepticism comes from yours.
In the end, depression is a social/psychological problem. It makes sense the the best cures would be too.
Excellent work. One friendly push on the arithmetic, because I think it may strengthen your conclusion rather than weaken it.
The 80/20 framing treats the drug effect and placebo response as separable line items: total improvement minus the placebo arm equals the “real” drug effect. But that implies additivity. Benedetti’s placebo work suggests something messier: expectancy and ritual recruit measurable endogenous opioid, dopaminergic, serotonergic, and stress/modulatory systems. As you suggest, context is a biological input rather than noise.
https://pubmed.ncbi.nlm.nih.gov/16280578/
The rTMS placebo literature points the same way (Razza et al. found a large sham response in depression trials that was associated with improvement in the active-treatment group, suggesting placebo response may be part of therapeutic response rather than merely a confound). https://pubmed.ncbi.nlm.nih.gov/29111404/
So the measured 2-point drug-placebo difference may not be a fixed pharmacologic payload sitting on top of an 8-point contextual effect as much as the drug’s marginal effect within an already powerful clinical context.
That matters for the lower-dose suggestion. Using lower doses to leverage placebo while reducing risk assumes dose and context are independent levers. If they interact, the trade may not hold.
I really liked your suggestion that clinicians can likely get more treatment benefit clinically by spending less time optimizing drug chemistry and more time on the "extra" context factors like trust, warmth, attention, and frequency of visits.
I don’t think the recommendation follows. The article shows placebo groups improve a lot, but that doesn’t show the improvement comes from warmth, follow-ups, hope-building, or behavioral activation.
It could just be belief in the pill.
Can you really derive the magnitude of the medication-only effect by subtracting the effect size of the placebo arm? Seems possible the effects of psychological expectation, social treatment and role, and pharmacology are not independently additive. You'd have to have something like a trial of people running placebo-controlled experiments at home to figure this out, but it seems very possible to me that the pharmacology alone could knock out some of the same 'depression points' that are covered by being attended to weekly by a kindly psychiatrist etc. I forget what this is called statistically but it's like the opposite of gestalt, the whole is less than the sum of its parts.
The study design you mention is the one I referenced by Leuchter, as being the best effort at this, but it unfortunately had big baseline differences in the exact thing they were trying to study.
The placebo response is agnostic about the source of improvement with placebo. We just don't have a great sense of which factors matter most, how they overlap, whether they are additive. Personally, I think baseline characteristics/natural history drive most outcomes, with expectancy and contact next.
What you seem to be getting at is the idea that prescribing an active medication, in clinical practice, would tap into other non-pharmacological factors like expectancy. Doing nothing, which is what placebo represents in the rest of medicine, wouldn't. And I totally agree with this. A great paper by Roose, Rutherford, and Thase makes a similar point to argue that the only way to generate real-world non-pharmacological factors is to prescribe a medication. They have presented some data that more frequent appointments has larger effects in the placebo than medication arm.
But I worry that your perspective makes the assumption that a prescription captures most of these non-pharmacological effects. I just don't think this assumption is fair to make or supported by evidence. I'm happy to be shown contrary evidence. This should be easy to prove with clinical trials of reduced contact and infrequent visits, and yet no one runs these trials.
My mental model isn't that the known factors might reduce to each other or be straightforwardly redundant, but rather that they might exert effects on the same nodes in a higher-order network that is currently undescribed. To make a toy example, if the factors delineated here (let's call them pharmacology, psychology, and social effects, or A, B, and C) all affect some three other unknown variables that are closer to the bone (call them 1, 2, and 3) it might be the case that A affects 1 and 2, and B affects 2 and 3, and C affects all three, but if 1 has already been pushed by A, you don't get further effect on 1 by also pushing C. So if you have a study arm that estimates the effects of B and C (like a placebo arm) and an arm that estimates the effects of A, B, and C (experimental arm) it wouldnt follow that the effect of A is approximately what's observed in the experimental arm - placebo arm.
Thank you for articulating this, as this was my initial reaction too. I won't deny the power of non-pharmacological/contextual effects on placebo-arm trial participants (speaking from personal experience participating in an ineffectively blinded study), but it strikes me as unlikely that it would be this simple.
I'm also skeptical that every single real-life patient who experiences meaningful improvement on an antidepressant derives the vast majority of that benefit from placebo response, especially if the average patient care experience is nothing more than a quarterly 10-minute check-in. Maybe the explanation is that responders receive an above-average level of care, but I don't know if that's been studied. Intuition of course doesn't always bear out in reality, but just offering my $0.02.
All fair points and my goal wasn't to suggest any of this is simple--I think it's really complicated, especially in practice.
Nils and I have argued about your point. This is where I feel like clinical trials offer the relevant counterfactual. There are just countless clinical trials, like the esketamine one I cite, that show massive improvement in the placebo arm.
The reality is that there is a range of response to both antidepressant and placebo, some people get a little better, some a lot better. But that's seen in both arms, like the figure from the Stone paper I showed.
Other hypotheses like increased variance among the antidepressant were disproven. https://jamanetwork.com/journals/jamapsychiatry/fullarticle/2776611
The parsimonious explanation is that non-pharmacological seem to matter most in common disorders--at least in the clinical trials that, ironically, lead to FDA med approval.