Skip to content

Student library · Evidence skills

How to Read a Randomised Trial

You are taught a checklist: randomised, concealed, blinded, intention-to-treat. The empirical question is which of those items actually changes the answer — and the honest reply is that it depends almost entirely on whether your outcome is subjective, and that in rehabilitation the pattern differs from the rest of medicine.

Evidence Concealment, subjective outcomes: ROR 0.69· Concealment, objective outcomes: 0.91 (ns)· Rehabilitation: 227,806 participants

In one line. A ratio of odds ratios below 1 means trials with that flaw reported bigger treatment effects than trials without it. That single statistic is how the field measures which design features matter, and it is the number this page is built around.

Read the design, then read the outcome, then decide how much the design mattered. Doing it in the other order is how a checklist becomes a ritual.

The single most useful thing on this page. With subjective outcomes — pain, function, patient-reported anything — inadequate or unclear allocation concealment exaggerated effects by a ratio of odds ratios of 0.69 (95% CI 0.59 to 0.82). With objective outcomes the same flaw produced 0.91 (0.80 to 1.03), which does not exclude no effect. [2] Physiotherapy outcomes are overwhelmingly subjective. Concealment is therefore the item to check first in the literature you will read.

What the evidence shows

Design featureWhat it does to the resultSource and quality
Allocation concealment, subjective outcomes Effects exaggerated. Ratio of odds ratios 0.69 (95% CI 0.59 to 0.82) when concealment was inadequate or unclear 146 meta-analyses, 1,346 trials, combining three meta-epidemiological studies [2]
Allocation concealment, objective outcomes Little evidence of bias. ROR 0.91 (0.80 to 1.03) Same study. The interval crosses 1 [2]
Blinding, subjective outcomes Effects exaggerated. ROR 0.75 (0.61 to 0.93) Same study. For objective outcomes, 1.01 (0.92 to 1.10) — nothing [2]
Blinding, re-examined in 2020 No evidence of an average difference. ROR 0.91 (0.61 to 1.34) for unblinded patients, 1.01 (0.84 to 1.19) for providers, 1.01 (0.86 to 1.18) for observers 142 meta-analyses, 1,153 trials. The authors say this may reflect that blinding matters less than believed, or study limitations, and recommend blinding remain a safeguard [4]
Sequence generation Effects exaggerated by an average of 11% when inadequate or unclear (ROR 0.89, 95% CrI 0.82 to 0.96), and worst for subjective outcomes (0.83, 0.74 to 0.94) 1,973 trials in 234 meta-analyses. Between-trial heterogeneity also rose, most of all for subjective outcomes [3]
Can assessors even agree on these judgements? Only moderately. Median kappa 0.60 for sequence generation, 0.58 for allocation concealment, 0.87 for blinding Same study. The item that matters most is the one two readers agree on least [3]
Sequence generation, binary outcomes High risk of bias made an exaggerated estimate far more likely: odds ratio 5.97 (95% CI 2.03 to 17.63). Incomplete outcome data: 4.07 (1.03 to 16.15) 101 studies, 128 pairwise comparisons. Note the very wide intervals [5]
And in rehabilitation specifically? Effects are likely exaggerated with inadequate or unclear sequence generation and allocation concealment when outcomes are continuous. The influence of blinding was inconsistent and different from the rest of medical science — overestimation, underestimation and neutral associations all appeared Cochrane Rehabilitation synthesis of 7 studies, 227,806 RCT participants. The pattern was more consistent for patient-reported outcomes [1]
Reporting bias in rehabilitation Appears to be associated with overestimation of treatment effects Same synthesis. Attrition and intention-to-treat were examined in only two studies, with inconsistent results [1]

Where students get this wrong

1. Applying the checklist without looking at the outcome

The same flaw is serious or irrelevant depending on what is being measured. Inadequate concealment moved subjective results by a ratio of odds ratios of 0.69 and objective results by 0.91 with an interval crossing 1. [2] "This trial was not blinded" is not a criticism until you say what it measured. A mortality trial and a pain trial are not equally threatened by the same defect.

2. Treating blinding as the master criterion

It is the item everybody remembers and its evidence is the least settled. The 2008 analysis found exaggeration with subjective outcomes; [2] the 2020 analysis of 142 meta-analyses found no average difference for patients, providers or observers, with intervals comfortably spanning 1. [4] In rehabilitation the direction was inconsistent and explicitly unlike the rest of medicine. [1] The defensible position is the one the 2020 authors took: keep blinding as a safeguard, stop treating it as the single marker of a trustworthy trial.

3. Ignoring allocation concealment because it sounds like randomisation

They are different. Sequence generation is how the list was made; concealment is whether the person recruiting could see what came next. Concealment is the item with the strongest and most consistent evidence of effect on subjective outcomes, [1][2] and it is also the item on which two assessors agreed least — median kappa 0.58. [3] It deserves more of your attention than it usually gets, and more care.

4. Reading "unclear" as "probably fine"

Every one of these analyses pools inadequate or unclear together, because trials that do not report a safeguard largely behave like trials that did not use one. An unreported method is a finding about the trial, not a gap in your reading.

5. Assuming the medical literature transfers to rehabilitation

It does for concealment and sequence generation. It does not for blinding, where the Cochrane Rehabilitation synthesis found the pattern inconsistent and different from the rest of medical science. [1] That group also concluded the evidence overall is mixed and inconclusive, partly because the methodological quality of rehabilitation trials is poor. Quoting a drug-trial rule of thumb at a physiotherapy trial is a mistake examiners increasingly notice.

6. Reading a wide confidence interval as a result

An odds ratio of 5.97 sounds decisive until you read the interval: 2.03 to 17.63. [5] It is genuinely significant and its magnitude is almost unknown. The same applies to 4.07 (1.03 to 16.15). Report the interval; the point estimate on its own is the number that gets misquoted.

What the evidence supports — and what it does not

Supported

  • Checking allocation concealment first when the outcome is subjective. [1][2]
  • Checking sequence generation. Average 11% exaggeration when inadequate or unclear. [3]
  • Treating "unclear" as a risk, not a neutral. [2][3]
  • Reading the outcome type before judging the design. [2]
  • Keeping blinding as a methodological safeguard even where its measured effect is uncertain. [4]

Not supported

  • Blinding as the decisive marker of trial quality. [4][1]
  • Applying drug-trial bias patterns to rehabilitation. [1]
  • Assuming design flaws bias objective outcomes. [2]
  • Firm conclusions about attrition or intention-to-treat in rehabilitation. Two studies, inconsistent. [1]
  • Quoting a point estimate without its interval. [5]

How certain is this?

Evidence grade: Moderate.

The design is unusually strong for a methodological question: these are meta-epidemiological studies, which pool many meta-analyses to ask whether a design feature is associated with a systematically different result. The samples are large — 1,346 trials, [2] 1,973 trials, [3] 1,153 trials, [4] and 227,806 participants in the rehabilitation synthesis. [1]

What holds the grade at moderate is that the field contradicts itself on blinding, [2][4] and that the rehabilitation-specific synthesis describes its own evidence as mixed and inconclusive, limited by the poor methodological quality of the underlying trials. [1] The concealment finding is the most secure; the blinding finding is the least.

A note on the dates. Three of these papers predate 2021. Meta- epidemiological studies are the primary derivations for these claims and do not expire the way treatment evidence does; where a newer analysis exists it is cited, and where it disagrees with the older one, both are reported.

Common questions

What should I check first in a trial?

Whether allocation was concealed, and what the primary outcome was. For subjective outcomes, inadequate or unclear concealment exaggerated effects by a ratio of odds ratios of 0.69; for objective outcomes it did not measurably bias them at all. [2] Almost every outcome you will read about in physiotherapy is subjective, so this one check does most of the work.

Does blinding matter or not?

Unresolved, and you should say so rather than pick a side. The 2008 combined analysis found exaggeration for subjective outcomes (ROR 0.75, 0.61 to 0.93); [2] the 2020 analysis of 142 meta-analyses found no average difference for patients, providers or observers. [4] In rehabilitation the direction was inconsistent. [1] Blinding remains a safeguard worth having; it is not proof of a trustworthy result and its absence is not proof of a biased one.

What is a ratio of odds ratios?

A comparison of comparisons. It asks: how much larger is the treatment effect in trials with a flaw than in trials without it? Below 1 means flawed trials reported bigger benefits. 0.69 means roughly a 31% exaggeration; 0.89 means about 11%. [2][3] An interval that includes 1 means no detectable difference.

Why do rehabilitation trials behave differently?

Partly because blinding a therapist is usually impossible, so unblinded trials in rehabilitation are not a self-selected group of sloppy studies the way they may be in pharmacology. The Cochrane Rehabilitation synthesis found the influence of blinding inconsistent and "different from the rest of medical science", though more consistent for patient-reported outcomes. [1] See also the PEDro scale, where the same constraint caps the score a good exercise trial can reach.

My appraisal disagrees with my partner's. Who is right?

Possibly neither, and the disagreement is normal. Median agreement between assessors was kappa 0.60 for sequence generation and 0.58 for allocation concealment — moderate at best. [3] Blinding was easier at 0.87. Resolve by going back to the paper's methods section together rather than by seniority; that is what consensus rating means, and it is why reviews use two assessors.

References

  1. Arienti C, Armijo-Olivo S, Ferriero G, et al. The influence of bias in randomized controlled trials on rehabilitation intervention effect estimates: what we have learned from meta-epidemiological studies. European Journal of Physical and Rehabilitation Medicine. 2024 Feb;60(1):135–144. doi:10.23736/S1973-9087.23.08310-7 PMID 38088137 Synthesis of meta-epidemiological studies
  2. Wood L, Egger M, Gluud LL, et al. Empirical evidence of bias in treatment effect estimates in controlled trials with different interventions and outcomes: meta-epidemiological study. BMJ. 2008 Mar 15;336(7644):601–5. doi:10.1136/bmj.39465.451748.AD PMID 18316340 Combined meta-epidemiological analysis
  3. Savović J, Jones H, Altman D, et al. Influence of reported study design characteristics on intervention effect estimates from randomised controlled trials: combined analysis of meta-epidemiological studies. Health Technology Assessment. 2012 Sep;16(35):1–82. doi:10.3310/hta16350 PMID 22989478 Meta-epidemiological study
  4. Moustgaard H, Clayton GL, Jones HE, et al. Impact of blinding on estimated treatment effects in randomised clinical trials: meta-epidemiological study. BMJ. 2020 Jan 21;368:l6802. doi:10.1136/bmj.l6802 PMID 31964641 Meta-epidemiological study
  5. Koletsi D, Spineli LM, Lempesi E, et al. Risk of bias and magnitude of effect in orthodontic randomized controlled trials: a meta-epidemiological review. European Journal of Orthodontics. 2016 Jun;38(3):308–12. doi:10.1093/ejo/cjv049 PMID 26174770 Meta-epidemiological study

About this resource

How to use this

Written to be learned from, not memorised.

This page is organised around which design flaws measurably change a result, rather than around the order a checklist happens to list them in. Faculty may use this page in teaching with attribution. It carries its review date and its next review date, so you can see at a glance whether it is current before you put it in front of a cohort.