Skip to content

Student library · Evidence skills

The PEDro Scale

You will be taught to score trials out of 10 and to treat 6 as the pass mark. The two largest validity studies of the scale disagree with each other about whether that total should be calculated at all — and one of them says plainly that it should not.

Evidence 11 items, 10 scored· Total-score reliability ICC 0.56 to 0.80· Agreement with Cochrane: kappa 0.12

In one line. The Physiotherapy Evidence Database scale rates the methodological quality of a randomised trial on 11 items, of which 10 count toward a score out of 10. Item 1, external validity, is deliberately not scored.

It exists because physiotherapy needed a quick, teachable way to sort trials. It is the right tool to learn first. It is not a tool to lean on, and this page is about the difference.

The exam answer and the honest answer are not the same, and you need both. For your exam: 11 items, 10 scored, higher is better, 6 or more is conventionally called moderate-to-high quality. For practice: a total score compresses items that measure different things into one number, and the study that tested whether that is legitimate concluded it is not. [2] Learn the scale, report the items.

What the evidence shows

QuestionFindingSource and quality
Can two people score the same trial the same way? Roughly. Item kappas ran 0.36 to 0.80 for individual raters and 0.50 to 0.79 for consensus ratings by two or three raters The scale's own reliability study. Described by its authors as "fair" to "substantial" [1]
How reliable is the total? ICC 0.56 (95% CI 0.47 to 0.65) for individual raters; 0.68 (0.57 to 0.76) for consensus ratings Same study. The authors call this "fair" to "good" — note that a single student scoring alone sits at the bottom of that range [1]
Is it more reliable elsewhere? Yes, in one setting: ICC 0.80 (0.68 to 0.88) for the total, with most individual items above 0.60 53 pharmaceutical trials of pain medication. Reliability figures do not automatically transfer between the literatures a tool is applied to [3]
Does the total measure one thing? No. There was substantial evidence of departure from the unidimensionality assumption, meaning the items relate to more than one underlying trait 345 trials from Cochrane reviews, item response theory. The authors conclude PEDro summary scores should not be used [2]
Which items actually carry the information? Allocation concealment and intention-to-treat analysis, with discriminations of 1.79 and 2.05. The other items provided little additional information and did not distinguish trials of different quality Same analysis. Mean summary score across the 345 trials was 5.46 (SD 1.51) out of 10 [2]
But is it valid at all? Yes, on a different test: correlation of 0.83 (95% CI 0.76 to 0.88) with the Cochrane Back and Neck risk of bias tool, inversely associated with treatment effect sizes, and not associated with journal impact factor This is the paper that disagrees with [2]. It tested convergent and construct validity rather than dimensionality, and in pharmaceutical rather than physiotherapy trials [3]
Does a PEDro cut-off pick the same trials as Cochrane? No. Agreement was poor at 5 or more (kappa 0.12, 95% CI 0.07 to 0.16) and slight at 6 or more (kappa 0.24, 0.16 to 0.32) 41 Cochrane reviews, 353 physiotherapy trials. Only 53.7% of the meta-analyses included trials of adequate quality by Cochrane criteria [4]
Does the choice change the answer? Sometimes. A substantial difference in the combined effect size (0.15 or more) appeared in 22% of meta-analyses at a cut-off of 5, and 24% at 6 Same study. In roughly one review in four, which quality tool you used changed the pooled result [4]
Why do physiotherapy trials score badly on blinding? Only 20% of 393 physiotherapy trials were adequately blinded, and most components of blinding were poorly reported 43 meta-analyses, 44,622 patients. Trials with inappropriate blinding tended to underestimate effects, but not significantly (assessors: effect size -0.07, 95% CI -0.22 to 0.08; participants: -0.12, -0.30 to 0.06) [5]

Where students get this wrong

1. Treating 6 out of 10 as a fact about the trial

The cut-off is a convention, not a property. Mean score across 345 trials in Cochrane reviews was 5.46, [2] so the average trial in the literature you are appraising sits just below the line everyone quotes. More importantly, two trials both scoring 6 can differ completely: one may have concealed allocation and analysed by intention to treat, the other may have scored its 6 on items the validity analysis found carried almost no information. [2] The number does not tell you which.

2. Assuming a low score means a bad trial

Three PEDro items concern blinding of participants, therapists and assessors. Blinding a therapist to whether they are delivering exercise is usually impossible. Only 20% of 393 physiotherapy trials were adequately blinded, [5] which is partly a statement about the trials and partly a statement about the field. A well-conducted exercise trial is structurally capped at a score that a mediocre drug trial clears easily. Read the items and ask which ones the design could realistically have met.

3. Assuming a low score means an inflated result

The intuition is that weak blinding exaggerates benefit. In physiotherapy the measured direction was the opposite, and not statistically significant: trials with inappropriate blinding of assessors and participants tended to underestimate treatment effects. [5] The authors are careful, and so should you be — the absence of significance is not evidence that blinding does not matter, only that this sample could not show it.

4. Quoting one validity paper and not the other

Two good studies reach opposite practical conclusions. The item response theory analysis of 345 trials says the items measure more than one trait and summary scores should not be used. [2] The validity study of 53 pharmaceutical trials reports strong convergence with the Cochrane tool (0.83) and total-score reliability of 0.80. [3] They are not directly comparable — different literatures, different questions — and this page does not pretend the disagreement is settled. If you cite PEDro's validity in an assignment, cite both.

5. Believing a quality tool is interchangeable with any other

PEDro and Cochrane agreed poorly on which trials were adequate: kappa 0.12 at a cut-off of 5. [4] That is barely above chance. In 22% to 24% of meta-analyses, switching tool changed the pooled effect size by 0.15 or more. [4] "Quality was assessed using the PEDro scale" is therefore a methodological choice with consequences, not a formality.

6. Scoring alone and trusting the result

Total-score reliability was ICC 0.56 for individual raters and rose to 0.68 for consensus between two or three. [1] If you are scoring trials for a dissertation, score with a partner and resolve differences; that is not a courtesy to your co-marker, it is the difference between "fair" and "good" reliability in the scale's own validation.

What the evidence supports — and what it does not

Supported

  • Using PEDro items to structure appraisal. Item reliability is acceptable for most items. [1][3]
  • Paying particular attention to allocation concealment and intention-to-treat. These carried most of the information. [2]
  • Scoring by consensus rather than alone. [1]
  • Reporting which tool you used and why. It changes results. [4]

Not supported

  • Reporting a PEDro total as a measure of trial quality. [2]
  • Treating 6 out of 10 as a validated threshold. [2][4]
  • Assuming PEDro and Cochrane classify trials alike. kappa 0.12. [4]
  • Reading a low blinding score as proof of an inflated effect. [5]
  • Transferring reliability figures from one literature to another. ICC 0.56 versus 0.80 in different samples. [1][3]

How certain is this?

Evidence grade: Moderate.

The individual studies are strong. The reliability study is the scale's own validation and reports its limits honestly. [1] The item response theory analysis uses 345 trials drawn from Cochrane reviews and applies a method designed to answer exactly the question it asks. [2] The meta-epidemiological work covers 353 and 393 trials respectively, the latter with 44,622 patients. [4][5]

What holds the grade at moderate is that the field has not resolved its own disagreement. [2] and [3] point in different directions, and no study has repeated the item response theory analysis in a larger or more recent sample. Until that happens, the defensible position for a student is the cautious one: use the items, be careful with the total.

A note on the dates. Four of these five papers were published before 2021, and that is deliberate. Methodological validation studies do not expire the way treatment evidence does, and these remain the primary derivations for the claims made about the scale. Where a newer study exists, this page cites it.

Common questions

My lecturer wants a PEDro total. What do I do?

Give them the total, and give them the items as well. The evidence does not support the total as a measure of quality, [2] but it is what the assignment asks for and it is what the PEDro database itself publishes. Add one sentence noting which items were met — particularly allocation concealment and intention-to-treat, which carried most of the information in the validity analysis. [2] That sentence is the difference between a pass and a good mark.

Why is item 1 not scored?

Item 1 concerns eligibility criteria, which is a question about external validity — who the results apply to — rather than internal validity, which is whether the result is likely to be true. The scale scores the remaining 10 items. That distinction is worth understanding, because a trial can be internally impeccable and still tell you nothing about your patient.

Is a trial scoring 4 out of 10 worthless?

No, and treating it that way will cost you marks. Ask which items it failed. If it lost three points on blinding in an exercise trial where blinding was impossible, that is a constraint of the field — only 20% of physiotherapy trials were adequately blinded. [5] If it lost points on allocation concealment and intention-to-treat, that is a genuine threat to the result. [2] Same total, different trial.

Should I use Cochrane's risk of bias tool instead?

For a systematic review, generally yes, and say which you used. The two tools agree poorly on which trials are adequate (kappa 0.12 at a cut-off of 5), [4] and in about a quarter of meta-analyses the choice changed the pooled effect by 0.15 or more. [4] PEDro is faster and better suited to rapid appraisal and to teaching; Cochrane is the standard for a review you intend to publish.

How does this connect to reading outcome measures?

The same discipline applies: a number is only interpretable with the study that produced it. The outcome measures library reports every reliability coefficient and threshold with its population and sample size for exactly this reason, and red flags shows what happens to a clinical rule when nobody checks the evidence behind it.

References

  1. Maher CG, Sherrington C, Herbert RD, et al. Reliability of the PEDro scale for rating quality of randomized controlled trials. Physical Therapy. 2003 Aug;83(8):713–21. PMID 12882612 Reliability study
  2. Albanese E, Bütikofer L, Armijo-Olivo S, et al. Construct validity of the Physiotherapy Evidence Database (PEDro) quality scale for randomized trials: Item response theory and factor analyses. Research Synthesis Methods. 2020 Mar;11(2):227–236. doi:10.1002/jrsm.1385 PMID 31733091 Item response theory and factor analysis
  3. Yamato TP, Maher C, Koes B, et al. The PEDro scale had acceptably high convergent validity, construct validity, and interrater reliability in evaluating methodological quality of pharmaceutical trials. Journal of Clinical Epidemiology. 2017 Jun;86:176–181. doi:10.1016/j.jclinepi.2017.03.002 PMID 28288916 Validity and reliability study
  4. Armijo-Olivo S, da Costa BR, Cummings GG, et al. PEDro or Cochrane to Assess the Quality of Clinical Trials? A Meta-Epidemiological Study. PLOS ONE. 2015;10(7):e0132634. doi:10.1371/journal.pone.0132634 PMID 26161653 Meta-epidemiological study
  5. Armijo-Olivo S, Fuentes J, da Costa BR, et al. Blinding in Physical Therapy Trials and Its Association with Treatment Effects: A Meta-epidemiological Study. American Journal of Physical Medicine & Rehabilitation. 2017 Jan;96(1):34–44. doi:10.1097/PHM.0000000000000521 PMID 27149591 Meta-epidemiological study

About this resource

How to use this

Written to be learned from, not memorised.

This page gives the exam answer and the honest answer side by side, because a student who can only give one of them is not yet able to appraise a trial. Faculty may use this page in teaching with attribution. It carries its review date and its next review date, so you can see at a glance whether it is current before you put it in front of a cohort.