Skip to content
Mindful Living
Programs
Entry F-033 · 4 minute read

Reading a Mindfulness Study

Four questions get a non-specialist most of the way through a paper in this field, and the first one settles most of it.

By Ruth Ellery Published

The literature on mindfulness is large and uneven, and most of the confusion in public coverage comes from a small number of design issues that are visible to any careful reader.

What was the control group

This is the question that settles most disagreements. A waitlist control receives nothing: no group, no weekly structure, no teacher, no expectation of improvement. Comparing an eight-week course against it measures the combined effect of all of those plus the meditation.

An active control holds the non-specific elements roughly constant. Effects measured against active controls are consistently smaller, and sometimes absent. When two studies appear to contradict each other, the comparator explains it more often than anything else.

How many people, and how many finished

Eight-week programmes with a daily assignment lose participants. Whether the analysis counts everyone who enrolled, or only those who completed, changes the result: completers are by definition the people who found it useful enough to continue.

Were the outcomes self-reported

Almost always, and for a programme aimed at a person’s relationship to their symptoms that is appropriate. It becomes a problem when the same evidence is used to argue for effects on disease processes rather than on experience.

Was the analysis planned in advance

Pre-registration is more common than it was and still not universal. An analysis chosen after the data arrived can find something in almost any dataset.

A note on brain imaging

Neuroimaging studies are the most reported and the most over-read. Samples are frequently under thirty, findings are at group level rather than per person, and a measurable change in a brain region is not evidence that anyone felt better. Treat imaging as evidence that something physical is happening, which was never in doubt, rather than as evidence of benefit.

Effect sizes, briefly

Reviews in this area typically report standardised effects around 0.3 to 0.6 for the outcomes that carry a moderate grade. That is a real and modest magnitude, broadly comparable to what is reported for structured exercise programmes or for antidepressant medication over placebo in milder presentations.

It is worth knowing that an average of this size is compatible with substantial benefit for some participants, none for others, and a small number who find the practice unhelpful. A trial reports the average; a person deciding whether to enrol is asking about themselves.

The attrition question

Eight-week programmes with a daily assignment lose people, and how a study handles that changes the headline. An analysis of completers only describes those who found it useful enough to continue, which flatters the result. An intention-to-treat analysis counts everyone who enrolled and produces smaller, more honest effects.

Where two studies seem to disagree, checking the analysis type resolves it surprisingly often.

Who gets studied

Participants in these trials skew female, well educated, and already sympathetic to the premise, because those are the people who volunteer. Whether the same effects appear in someone sceptical and referred rather than self-selected is genuinely open, and it is the question a routine clinical referral is implicitly asking.

Common questions

Where can these papers be read for free?

PubMed Central hosts free full text for a large share of the clinical literature in this area, including much of the frequently cited work. Systematic reviews are the efficient entry point because they grade the underlying trials for you.

Is a meta-analysis always better than a single trial?

Usually, because it grades consistency across studies rather than reporting one result. It inherits the weaknesses of what it pools, though, so a meta-analysis of small waitlist-controlled trials is not strong evidence however many trials it contains.

What counts as a good sample size here?

It depends on the effect being looked for, and as a rough guide neuroimaging studies under thirty participants should be read cautiously, while clinical trials in this field commonly run to a few hundred.

Why does the same intervention produce different results in different studies?

Most often the comparator differs. Beyond that: different populations, different teacher quality, different adherence to home practice, and different outcome measures for what sounds like the same thing.