Skip to content

Student library · Evidence skills

GRADE

GRADE looks like a calculation and is a structured judgement. When experienced reviewers applied it to the same evidence, agreement on the overall grade was slight — and the authors concluded that their conclusions "can differ greatly".

Evidence Overall grade agreement: slight· Risk-of-bias domains: kappa 0.24 to 0.37· Rated per outcome, not per study

In one line. GRADE rates how much confidence you can place in an estimate of effect, as high, moderate, low or very low. Randomised trials start high and are rated down for risk of bias, inconsistency, indirectness, imprecision and publication bias; observational studies start low and can be rated up.

The single most important structural fact, and the one most often missed in assignments: GRADE is applied per outcome, not per study and not per review. The same review can hold high-certainty evidence for pain and very-low for function.

Low certainty does not mean the treatment does not work

It means the estimate could change substantially with further research. "Low-certainty evidence that TENS reduces pain" is a statement about how much the number might move, not a verdict on the treatment. Writing "the evidence shows TENS does not work" when the review says low certainty is the single most common error in student appraisals, and it inverts the meaning.

What the evidence shows

QuestionFindingSource and quality
Do experienced reviewers reach the same overall grade? Largely no. Agreement on independent reviewers' strength-of-evidence grades was generally poorer than for the individual domains, and overall agreement was slight — not appreciably better even when restricted to exercises containing only randomised trials Formal reliability testing of the AHRQ grading approach. The authors conclude that conclusions reached by experienced reviewers from the same evidence "can differ greatly" [1]
Does agreeing on the domains fix it? No. Neither agreement on domain scores nor agreement about how difficult a domain was to evaluate predicted the overall grade Same study. Domain agreement itself ranged from substantial (risk of bias in RCTs, directness) down to slight (risk of bias in observational studies) [1]
How reliable is the risk-of-bias input? Fair. Between two reviewers, kappa ranged 0.24 to 0.37 for most Cochrane risk-of-bias domains, with sequence generation the exception at 0.79 154 randomised trials. Consensus assessments across four reviewer pairs were worse still: 0.60 for sequence generation, 0.37 for allocation concealment, and 0.05 to 0.09 for the remaining domains [2]
Why do reviewers disagree? Mostly about the tool, not the trial. Inter-rater variability resulted more often from different interpretation of the instrument than from other causes, and was influenced by the nature of the outcome, the intervention, the design, the hypothesis and the funding source Same report [2]
Can a structured aid help? Partly. Two independent researchers using a structured assessment tool agreed well with each other (overall rating weighted kappa 0.79, 95% CI 0.65 to 0.93), but the tool agreed only fair-to-moderately with a standard GRADE assessment (weighted kappa 0.35, 95% CI 0.09 to 0.87) Note the width of that second interval — 0.09 to 0.87 carries very little information. The authors suggest it may aid consistency, particularly for less experienced researchers [3]
What does downgrading look like in practice? TENS versus placebo: pain lower during or immediately after treatment (SMD -0.96, 95% CI -1.14 to -0.78), moderate certainty. TENS versus standard care: SMD -0.72 (-0.95 to -0.50), low certainty — downgraded because small trials made the magnitude imprecise 381 randomised trials, 24,532 participants. The same review, the same intervention, two different certainty ratings for two different comparisons [4]
How often is certainty high? Rarely. Across 8 reviews of shockwave therapy for knee osteoarthritis, of 49 GRADE-assessed outcome indicators only 3 were moderate and the rest were low or very low. None were high Limitations — the risk-of-bias domain — were the commonest reason for downgrading [5]

Where students get this wrong

1. Reading "low certainty" as "ineffective"

The four levels describe confidence in the estimate, not the direction of the effect. The TENS review reports a statistically clear reduction in pain against standard care (SMD -0.72, 95% CI -0.95 to -0.50) at low certainty, [4] which means the effect is probably real and the size of it may change with better trials. If you write that low certainty means the treatment failed, you have reported the opposite of the finding.

2. Grading the review instead of the outcome

GRADE attaches to an outcome within a comparison. The TENS review carries moderate certainty for one comparison and low for another in the same paper. [4] A sentence beginning "this review was graded low" is a category error; the correct form names the outcome and the comparison.

3. Treating it as arithmetic

Two experienced reviewers working from the same evidence produced overall grades that agreed only slightly, and agreement on the underlying domains did not predict the final grade. [1] There is no formula that converts five domain judgements into a level. That is a design feature of GRADE, not a flaw in your understanding, and the published guidance expects the reasoning to be written down for exactly this reason.

4. Trusting the risk-of-bias input more than it deserves

GRADE's first downgrade domain rests on a risk-of-bias judgement whose inter-rater kappa ran 0.24 to 0.37 across most domains, and whose consensus agreement across four reviewer pairs fell to 0.05 to 0.09 for most domains. [2] The variability came mainly from different interpretations of the tool. When you disagree with a published rating, you are often observing normal variation rather than an error.

5. Downgrading twice for the same problem

Small trials cause imprecision; small trials are also often at higher risk of bias. It is tempting to take a level off for each. The TENS review is instructive here: the downgrade to low certainty was attributed specifically to small-sized trials contributing to imprecision in magnitude estimates, [4] a single named reason. Name the domain you are downgrading for, once.

6. Expecting high certainty in rehabilitation

You will rarely see it. Of 49 outcomes in one overview, three reached moderate and none reached high. [5] That is partly the state of the trials and partly structural — blinding constraints in physiotherapy push the risk-of-bias domain down before anything else is considered. See the PEDro scale for the same ceiling in a different tool.

What the evidence supports — and what it does not

Supported

  • Using GRADE to structure and record a certainty judgement. [1][3]
  • Rating per outcome and per comparison. [4]
  • Writing down the reason for every downgrade. [4][5]
  • Using a structured aid if you are inexperienced. Agreement between two users was 0.79. [3]
  • Expecting low certainty in rehabilitation and saying so plainly. [5]

Not supported

  • Treating a GRADE level as reproducible between assessors. Overall agreement was slight. [1]
  • Reading low certainty as evidence of no effect. [4]
  • Applying one grade to a whole review. [4]
  • Assuming domain agreement produces grade agreement. It did not. [1]
  • Treating published risk-of-bias ratings as definitive. kappa 0.24 to 0.37. [2]

How certain is this?

Evidence grade: Moderate.

The reliability findings come from formal testing programmes commissioned to answer exactly this question, using multiple reviewer pairs across several centres, [1][2] which is a stronger design than the usual two-rater study. The two reports agree with each other and with the smaller agreement study. [3]

What holds the grade at moderate is scope. These are assessments of the AHRQ grading approach and the Cochrane risk-of-bias tool rather than of GRADE exactly as a Cochrane author applies it today, and the tools have been revised since. The direction — that overall certainty judgements are considerably less reproducible than the confident four-level output suggests — is secure. The precise coefficients would move under current versions.

A note on the dates. Three of these sources predate 2021. They are the primary reliability derivations for the framework being taught, and reliability studies do not expire the way treatment evidence does. The two worked examples are recent.

Common questions

What do the four levels actually mean?

They describe how likely further research is to change the estimate. High: further research is very unlikely to change confidence in the effect. Moderate: likely to have an important impact and may change the estimate. Low: very likely to have an important impact and likely to change the estimate. Very low: any estimate is very uncertain. All four are statements about the estimate, not about whether the treatment works.

My grade differs from the published one. Am I wrong?

Possibly not. Agreement between experienced reviewers on overall grades was slight, and the authors state plainly that conclusions from the same evidence can differ greatly. [1] What matters in an assignment is not matching the published grade but naming the domain you downgraded for and the reason. Disagreement with your reasoning shown is a good answer; agreement with no reasoning is not.

Why do randomised trials start at high and observational studies at low?

Because randomisation addresses confounding by design, and observational designs cannot. Observational evidence can be rated up for a large effect, a dose-response gradient, or where plausible confounding would work against the observed effect. In rehabilitation you will mostly be rating trials down rather than observational studies up.

How does GRADE relate to AMSTAR 2?

They answer different questions and neither substitutes for the other. AMSTAR 2 rates how well a systematic review was conducted; GRADE rates how certain the evidence for a given outcome is. A comparison of the leading review-appraisal tools found that neither AMSTAR 2 nor ROBIS addresses certainty of evidence at all — see appraising a systematic review.

Is it worth using a structured worksheet?

If you are learning, yes. Two independent researchers using a structured assessment tool agreed at a weighted kappa of 0.79, and its authors suggest it may aid consistent application particularly for less experienced researchers. [3] Be aware that the same study found only fair-to-moderate agreement between the tool and a standard GRADE assessment (0.35, 95% CI 0.09 to 0.87), so a worksheet makes you consistent with yourself rather than correct.

References

  1. Berkman ND, Lohr KN, Morgan LC, et al. Reliability Testing of the AHRQ EPC Approach to Grading the Strength of Evidence in Comparative Effectiveness Reviews. AHRQ Methods for Effective Health Care. 2012 May. PMID 22764383 Reliability testing report
  2. Hartling L, Hamm M, Milne A, et al. Validity and Inter-Rater Reliability Testing of Quality Assessment Instruments. AHRQ Methods for Effective Health Care. 2012 Mar. PMID 22536612 Reliability and validity testing report
  3. Llewellyn A, Whittington C, Stewart G, et al. The Use of Bayesian Networks to Assess the Quality of Evidence from Research Synthesis: 2. Inter-Rater Reliability and Comparison with Standard GRADE Assessment. PLOS ONE. 2015;10(12):e0123511. doi:10.1371/journal.pone.0123511 PMID 26716874 Instrument development and agreement study
  4. Johnson MI, Paley CA, Jones G, et al. Efficacy and safety of transcutaneous electrical nerve stimulation (TENS) for acute and chronic pain in adults: a systematic review and meta-analysis of 381 studies (the meta-TENS study). BMJ Open. 2022 Feb 10;12(2):e051073. doi:10.1136/bmjopen-2021-051073 PMID 35144946 Systematic review and meta-analysis
  5. Zhou Q, Chen J. A critical overview of systematic reviews and meta-analyses of extracorporeal shockwave therapy for knee osteoarthritis. Asian Journal of Surgery. 2024 Jul;47(7):2975–2984. doi:10.1016/j.asjsur.2024.01.127 PMID 38290944 Overview of systematic reviews

About this resource

How to use this

Written to be learned from, not memorised.

This page treats GRADE as what it is, a structured judgement with measured limits on its reproducibility, rather than as a calculation with a right answer. Faculty may use this page in teaching with attribution. It carries its review date and its next review date, so you can see at a glance whether it is current before you put it in front of a cohort.