Reading a Mindfulness Study
Four questions get a non-specialist most of the way through a paper in this field, and the first one settles most of it.
The literature on mindfulness is large and uneven, and most of the confusion in public coverage comes from a small number of design issues that are visible to any careful reader.
What was the control group
This is the question that settles most disagreements. A waitlist control receives nothing: no group, no weekly structure, no teacher, no expectation of improvement. Comparing an eight-week course against it measures the combined effect of all of those plus the meditation.
An active control holds the non-specific elements roughly constant. Effects measured against active controls are consistently smaller, and sometimes absent. When two studies appear to contradict each other, the comparator explains it more often than anything else.
How many people, and how many finished
Eight-week programmes with a daily assignment lose participants. Whether the analysis counts everyone who enrolled, or only those who completed, changes the result: completers are by definition the people who found it useful enough to continue.
Were the outcomes self-reported
Almost always, and for a programme aimed at a person’s relationship to their symptoms that is appropriate. It becomes a problem when the same evidence is used to argue for effects on disease processes rather than on experience.
Was the analysis planned in advance
Pre-registration is more common than it was and still not universal. An analysis chosen after the data arrived can find something in almost any dataset.
A note on brain imaging
Neuroimaging studies are the most reported and the most over-read. Samples are frequently under thirty, findings are at group level rather than per person, and a measurable change in a brain region is not evidence that anyone felt better. Treat imaging as evidence that something physical is happening, which was never in doubt, rather than as evidence of benefit.
Effect sizes, briefly
Reviews in this area typically report standardised effects around 0.3 to 0.6 for the outcomes that carry a moderate grade. That is a real and modest magnitude, broadly comparable to what is reported for structured exercise programmes or for antidepressant medication over placebo in milder presentations.
It is worth knowing that an average of this size is compatible with substantial benefit for some participants, none for others, and a small number who find the practice unhelpful. A trial reports the average; a person deciding whether to enrol is asking about themselves.
The attrition question
Eight-week programmes with a daily assignment lose people, and how a study handles that changes the headline. An analysis of completers only describes those who found it useful enough to continue, which flatters the result. An intention-to-treat analysis counts everyone who enrolled and produces smaller, more honest effects.
Where two studies seem to disagree, checking the analysis type resolves it surprisingly often.
Who gets studied
Participants in these trials skew female, well educated, and already sympathetic to the premise, because those are the people who volunteer. Whether the same effects appear in someone sceptical and referred rather than self-selected is genuinely open, and it is the question a routine clinical referral is implicitly asking.