We Throw Out the Patients Who Get Better
Before a great many antidepressant trials formally begin, there is a phase that never makes it into the press release. Everyone who enrolls is given a sugar pill for a week or two. Nobody is told. The people who improve are removed from the study.
This is called a placebo run-in. It is not fringe methodology. It is standard, respectable, and reviewers ask for it, because in psychiatric trials the response to a sugar pill is enormous and the gap between the drug and the sugar pill is small enough to drown in it.
When a run-in is not enough, there is a second technique with a bureaucratic name: the Sequential Parallel Comparison Design. Stage one, split the cohort, drug and placebo. A large share of the placebo group improves, as expected. Stage two, take only the people who did not improve on a placebo, and run the trial again on them. You now have a study population selected specifically for being unable to benefit from being cared for.
Sit with the shape of that. We have built a multi-billion dollar research apparatus whose first operation is to find the people who heal inside a context of care and remove them, so that we can see the molecule.
Midas’s Gold Standard
In one of the larger pooled analyses of antidepressant trials, covering more than 200 studies and over seventy thousand patients, the average improvement on the seventeen-item Hamilton was about 10.1 points on active drug and 8.3 points on placebo. The difference is 1.8 points. Standardized, that is an effect size somewhere near 0.30. The usual threshold for a difference a patient could actually perceive is around three Hamilton points, or an effect size of 0.50.
So the drugs beat placebo, reliably, and do not beat it by enough for the person taking it to feel the gap.
Then there is the matter of active placebos. Sugar pills have no side effects. Real psychiatric medications produce dry mouth, mild nausea, a little dizziness. Patients notice, correctly conclude they got the real thing, and their expectations rise accordingly. When trials use an active placebo, something inert that mimics the side effect profile, the margin tends to shrink further toward nothing. Which means some real portion of what we have been calling drug effect is not chemistry at all. It is a patient successfully guessing their assignment.
Four Things the Method Requires
None of this is a scandal. It is a mismatch, which is worse, because scandals get corrected and mismatches get institutionalized.
The randomized controlled trial is a beautiful instrument built for a specific kind of object. It needs four things. An intervention that can be separated from its delivery. A comparator that is genuinely inert. A blind that holds. An outcome measurable independent of the patient’s own report.
Psychotherapy fails all four, and not marginally.
You cannot blind a person to whether they spent fifty minutes with someone who listened to them. You cannot blind the clinician. And the control condition, supportive attention, an empathic person working from a manual that forbids the specific technique, is not inert. It is the active ingredient with the branding stripped off. When your placebo contains the drug, a null result tells you nothing, and a positive result tells you only the marginal value of a technique layered on top of the thing that was already working.
We have measured this directly. Ted Kaptchuk’s irritable bowel study divided patients three ways: a waitlist, sham acupuncture delivered briskly by a practitioner who said almost nothing, and the identical sham acupuncture delivered by a warm practitioner who took time, asked questions, and projected confidence. The waitlist barely moved. The cold ritual produced moderate relief. The warm ritual produced relief that competed with the best drugs on the market for that condition. Same fake needles. The variable was the human being.
That is a dose-response curve for attention. In any other corner of medicine we would call it a finding and begin optimizing. In ours we call it a confound and begin controlling for it. Modern trials now use centralized remote raters, independent clinicians scoring symptoms over video who have never met the patient, specifically so that a person who has come to know a patient cannot accidentally leak warmth into the measurement.
The Finding That Should Have Ended the Argument
Open-label placebos.
You hand the patient a bottle. You tell them plainly these are sugar pills, nothing in them. You explain that the placebo response is real, that it appears to work partly through conditioning, that believing in it helps but is not required, and that taking them faithfully matters. Then you let them go.
They work. Across clinical populations, meta-analyses put the effect against no treatment in the moderate to large range. There are pilot randomized trials in major depressive disorder showing meaningful movement. The deception, which was supposed to be the load-bearing wall of the whole phenomenon, turns out not to be load-bearing.
The imaging is stranger still. When people take an open-label placebo for emotional distress, activation shows up in the periaqueductal gray and the hippocampus. There is no prefrontal activation, and the response does not track conscious expectation. The part of the brain that knows the pill is fake is not the part that responds. The ritual is talking to something older than the argument, and it does not appear to be listening to the argument at all.
As a therapist who watches patients get better in somatic and experiential therapy, often after they have failed to make progress in mannualized cognitive and behavioral models for years, I feel this same insecurity. It is a normal and an important part on reflecting on what does work. I often wonder did I do anything, or is this real? However years into therapy I still see the non quantifiable and self evidencing parts of the pattern of healing pay off again and again. I believe the patient when they tell me knew insights and new freedoms they are discovering as a result of the process. I am not insecure enough to write these off as placebo just because I cannot count them or distill their mechanism to one variable.
Now the Mechanism Problem
Psychology has spent forty years trying to buy a seat at the medical table with mechanism stories. Bilateral stimulation does this to memory reconsolidation. Cognitive restructuring does that to the schema. Every modality arrives with a brain diagram and an arrow.
There are two problems with this, and only one of them is scientific.
The first is that the mechanism story is doing commercial work. A mechanism is what you need for a training institute, a certification pathway, a patent, a grant renewal. I am not making an accusation, and I am not exempt. I have a mechanism story of my own that I have worked on for years and believe in, and I have caught myself loving the story more than I love the evidence for it. That is the whole trick of an incentive structure. It never requires anyone to lie. It only requires that the version of the truth that pays be slightly easier to keep believing.
The second problem is that in the drug literature, mechanism logic has already collapsed on its own terms.
The emerging genetics of placebo response, what researchers have started calling the placebome, keeps landing on the same systems the drugs target. COMT and prefrontal dopamine clearance. OPRM1 and the mu-opioid receptor. Serotonin transporter variants. Meanwhile the entire trial design rests on an additivity assumption: total effect equals drug effect plus placebo effect, so subtract the second to reveal the first. If both are running through the same receptors, they are not additive. They are competing for the same seats. The subtraction at the heart of the method is arithmetic performed on quantities that are not independent, and nobody has a good answer for this.
There is one more finding in that literature that should have been on the front page. Depressed patients who mount a strong placebo response, measured as endogenous opioid release in the nucleus accumbens and subgenual cingulate, go on to respond better to the actual antidepressant. That placebo-induced release predicted a substantial share of the variance in eventual drug response.
Read that again against the run-in period. The capacity to respond to care is not gullibility and it is not weak pathology. It is a functioning healing system. We have spent thirty years engineering it out of our trials and calling the removal rigor.
The Reductio Arrived and We Ignored It
The psychedelic trials are where this becomes impossible to look away from. You cannot blind a person to psilocybin. Everyone in the room knows, including the rater. And the outcome does not fit the instrument. People come back describing a reconciliation with death, a dissolution of the self, a wholesale reordering of what their life was for, and we hand them a form and ask them to rate, zero to four, their loss of interest in usual activities.
You cannot weigh a resurrection on a scale that counts symptoms.
The most promising development in psychiatry in fifty years is the one that most clearly breaks the measuring device, and the field’s response has been to build a more elaborate blind rather than a better instrument.
Rigor Is Not the Same Thing as Ritual
I want to be careful, because this argument gets misread as anti-science by people with an interest in misreading it. I am not against evidence. I am against ceremony.
A real science fits its method to the world. A ritual cuts the world down to fit the method and calls the cutting objectivity.
Gödel’s point, in the plain version, is that no system can prove itself from the inside. At the bottom of every body of knowledge there is something taken as given that cannot be established from within. Polanyi’s version is friendlier: we know more than we can tell. Everything explicit and writable floats on a much larger body of tacit knowledge, the thing an experienced clinician registers in the first ninety seconds and could not get onto a worksheet if you gave her a year. When we announce that what cannot be measured is not real, we believe we are being stricter. We are being smaller. We kept the lit tip of the iceberg and sawed off everything holding it up.
Hayek titled his Nobel lecture The Pretence of Knowledge and used the microphone to warn that imitating the methods of physics in a field where they do not fit forces you to discard the scattered, unmeasurable, lived knowledge that actually runs a complex system. We gave him the medal and built the next fifty years on the pretence.
Here is the shape of the failure, and it is the same shape at every scale. You can run a company entirely from the dashboard. Targets green, quarterly immaculate, books balanced, right up until somebody walks into the actual warehouse and finds it empty. The inventory system was perfect the entire time. The number was never the warehouse. It was a description of the warehouse, and a description can stay flawless long after the thing it describes has been hollowed out, because at some point it stopped being attached to the thing and started being attached to itself.
We run patients this way now. The chart looks treated. The scores trend down. The measure has come loose from the person and nothing inside the system can detect the drift, because the only instrument we kept is the one that drifted.
What Actually Replaces It
Not nothing, and not vibes. The answer is not less rigor. It is rigor aimed at the correct object.
Keep the randomized trial where the ontology fits. A micronutrient. A compound with a blindable side effect profile. A procedure against no procedure. The method is not the problem. The universal application of it is.
For everything relational, we already have better tools and we have made them low status.
Routine outcome monitoring, which for actual clinical utility outperforms most trial findings, because it tells you about this patient rather than about a hypothetical average one who does not exist.
Single-case experimental designs and n-of-1 trials, which are rigorous, replicable, publishable, and match the unit that actually walks into the office.
Practice research networks and naturalistic cohorts. Effectiveness instead of efficacy, in the conditions where the treatment will really be delivered.
Studying therapist effects instead of controlling for them. Therapist variance is among the most reliable findings in the entire psychotherapy literature, and the standard design treats it as noise to be averaged away. We are averaging away the signal and publishing the residue.
Qualitative work and the case study restored to standing rather than merely tolerated. The case study built most of what we know. We exiled it for looking insufficiently like chemistry.
And finally, stop demanding a mechanism before you will permit yourself to notice that something works. Aspirin was in use for roughly seventy years before anyone could explain it. Lithium reorganized psychiatry and we still cannot fully account for it. We put patients under general anesthesia in this country every single day without a settled explanation of how consciousness departs and returns. Medicine has never actually required mechanism. It required results, replication, and a plausible safety story. Psychology demands mechanism because psychology is insecure, and the demand has quietly become a tool for disqualifying things that work.
Why This Is Your Problem Too
If you are a physician reading this and filing it under someone else’s specialty, look at your own fifteen minutes.
The thing being stripped out of your visit is the identical thing our trials are engineered to remove. Not because anyone decided to remove it, but because it cannot be patented, cannot be billed at a rate that reflects what it does, cannot be scaled, and cannot be demonstrated by the one method we have agreed to accept as proof. A regulator can approve a molecule. There is no pathway for approving fifteen unhurried minutes and a person who believes you.
The relationship was never the packaging around the intervention. In an enormous amount of medicine, it is the intervention.
Two people in a quiet room. It is still the only thing that has ever consistently worked, and it is the exact variable our best instrument was specifically designed to control for.



























0 Comments