Student library · Evidence skills
Appraising a Systematic Review
A systematic review sits at the top of the pyramid you were shown in first year. It is also the design where two competent reviews of the same question routinely reach different conclusions, and where the standard appraisal tools agree with each other only moderately.
In one line. AMSTAR 2 rates the methodological quality of the review. It does not rate the certainty of the evidence the review contains, and it does not tell you what the treatment effect was. Those are three separate questions and students routinely collapse them into one.
Learn to say which of the three you are answering. That sentence alone will distinguish your critical appraisal from most of the ones your marker reads.
The gap nobody mentions. A comparison of the two leading tools found that neither AMSTAR-2 nor ROBIS addresses the reporting of absolute estimates of effect or the overall certainty of the evidence — items the authors call crucial for anyone wanting to interpret how important the results are. [3] A review can score highly on AMSTAR 2 and still contain evidence graded very low. That is not a contradiction; they measure different things.
What the evidence shows
| Question | Finding | Source and quality |
|---|---|---|
| Do two people rate a review the same way? | Moderately. Median agreement across all AMSTAR-2 questions was kappa 0.61, with 8 of the 16 questions reaching substantial agreement or higher | 31 systematic reviews. ROBIS scored identically overall (median 0.61) with 11 of its 24 questions substantial or higher [1] |
| Is the total more reliable than the items? | Yes. Median item-level kappa 0.60 (interquartile range 0.36 to 0.71), but the inter-observer ICC for the total score was 0.85 (95% CI 0.74 to 0.92) | 30 reviews assessed with AMSTAR. Note the interquartile range: some items were rated at kappa 0.36 [2] |
| Which items do reviews actually fail? | Reported in a third of reviews or fewer: an a priori protocol, the full list of included and excluded studies, quality assessment of included studies, publication bias assessment, and conflicts of interest | Same study [2]. An overview of shockwave therapy reviews found the same failures: no reasons for selection, no excluded-studies list, no reporting-bias assessment, no conflict of interest [5] |
| Do the two main tools measure the same thing? | Mostly. 70.3% of items (26 of 37) address the same or similar methodological constructs, with agreement almost perfect on six comparisons, substantial on three and moderate on one | Each tool also covers ground the other does not — AMSTAR-2 uniquely addresses funding of primary studies and reviewers' conflicts of interest [3] |
| What do neither of them cover? | Absolute estimates of effect, and the overall certainty of the evidence | Explicitly flagged by the comparison study as crucial and missing from both [3] |
| Two reviews disagree. How do I choose? | Not with the published algorithm. In 62% of cases (13 of 21) researchers could not replicate a Jadad decision and ended up choosing a different review | The algorithm has no prescriptive instructions for operationalising it. Notably, 86% (18 of 21) reached the same direction of finding despite picking a different review [4] |
| So what should you do instead? | Choose the review with meta-analysis of randomised trials that most closely resembles your clinical question, is most recent, most comprehensive by number of included trials, and at lowest risk of bias | The authors' own recommendation, offered in the absence of a validated algorithm [4] |
| What does this look like in practice? | Across 8 reviews of shockwave therapy for knee osteoarthritis, methodological quality was "generally unsatisfactory", and of 49 GRADE-assessed outcome indicators only 3 were moderate quality — the rest low or very low, with limitations the commonest reason for downgrading | A worked example of the whole problem: a positive conclusion resting on evidence the same authors graded as low [5] |
Where students get this wrong
1. Treating "systematic review" as a quality claim
It is a description of method, not a guarantee of rigour. In one overview, eight reviews of the same intervention were rated generally unsatisfactory, and the evidence inside them graded low or very low for 46 of 49 outcomes. [5] The label at the top of the pyramid tells you how the studies were gathered, not how good they were.
2. Confusing review quality with evidence certainty
These are separate assessments with separate tools, and the appraisal tools deliberately do not cross over — neither AMSTAR-2 nor ROBIS rates certainty of evidence or reports absolute effects. [3] A methodologically excellent review of four small biased trials is still a report about four small biased trials. Say which you are appraising.
3. Reporting an AMSTAR 2 rating as if it were exact
Median item agreement was kappa 0.61 in one study [1] and 0.60 with an interquartile range down to 0.36 in another. [2] Half the AMSTAR-2 questions did not reach substantial agreement between raters. [1] The total is more stable than the items (ICC 0.85), [2] so if you must give one figure, give the overall rating and note that item-level judgements vary between assessors.
4. Assuming reviews that disagree cannot both be sound
They often can. Differences in inclusion dates, eligibility criteria, and which outcome was chosen as primary will move a pooled estimate without anyone doing anything wrong. The useful finding is that when researchers picked different reviews from the same set, 86% still agreed on the direction of the answer. [4] Disagreement about magnitude is common; disagreement about direction is the one worth investigating.
5. Looking for a rule to pick the winner
There is no validated algorithm. The Jadad approach was not reproducible between users in 62% of cases because it never specified how to apply it. [4] What replaced it is judgement with stated criteria: closest to your question, most recent, most comprehensive, lowest risk of bias. Write those four down in your appraisal and you have made a defensible choice rather than an arbitrary one.
6. Skipping the boring items
The items reviews most often fail are the unglamorous ones: a pre-registered protocol, the list of excluded studies, publication bias, conflicts of interest. [2][5] They are also the ones that tell you whether the review's authors decided what to look for before or after they saw the results. That is the question the whole tool exists to answer.
What the evidence supports — and what it does not
Supported
- Using AMSTAR 2 or ROBIS to appraise review methodology. Both perform comparably. [1][3]
- Reporting the overall rating rather than item scores where one figure is needed. [2]
- Rating certainty separately, because neither tool does it. [3]
- Choosing between discordant reviews on stated criteria. [4]
- Checking protocol registration, excluded studies, publication bias and conflicts of interest first. [2][5]
Not supported
- Treating an AMSTAR 2 rating as a statement about the evidence. [3]
- Treating item-level ratings as reproducible. Half fell below substantial agreement. [1]
- Using the Jadad algorithm to resolve discordance. Not reproducible in 62% of cases. [4]
- Assuming a review's positive conclusion reflects certain evidence. 3 of 49 outcomes were moderate; none were high. [5]
- Assuming AMSTAR-2 and ROBIS are interchangeable. 70.3% overlap leaves real differences. [3]
How certain is this?
Evidence grade: Moderate.
The reliability findings are consistent across independent studies in different fields, which is a point in their favour: median agreement of 0.61 in 31 human-health reviews [1] and 0.60 in 30 veterinary reviews [2] is the kind of replication that makes a methodological finding believable.
What limits the grade is sample size. These are studies of 21 to 31 reviews, [1][4] and the reproducibility finding on discordance rests on 21 cases. [4] The direction is clear; the precise agreement coefficients would move with a larger sample. The worked example is a single overview in a single condition. [5]
A note on the dates. Three of these papers predate 2021. They are the primary reliability derivations for the tools you are being taught, and reliability studies do not expire the way treatment evidence does.
Common questions
AMSTAR 2 or ROBIS?
Either. Median inter-rater agreement was identical at 0.61, and 70.3% of their items address the same constructs. [1][3] AMSTAR-2 was described as more straightforward to use; ROBIS was easier to apply to reviews containing a meta-analysis. [1] Use whichever your course specifies, name it, and do not switch tools halfway through an assignment.
My review scored "critically low". Is it useless?
No, but say what drove the rating. The commonest failures are protocol registration, the excluded-studies list, publication bias and conflicts of interest, [2][5] and a review can fail several of those and still have searched well and pooled correctly. Report which items failed rather than the label alone — that is the same argument as on the PEDro scale page, and for the same reason.
How do I appraise certainty, if not with AMSTAR?
With GRADE, which is a separate assessment applied per outcome rather than per review. The shockwave overview shows why it matters: 49 outcome indicators, 3 moderate, the rest low or very low, with limitations the commonest downgrade reason. [5] Neither AMSTAR-2 nor ROBIS produces that judgement. [3]
Two reviews of the same question say different things. What do I write?
Write the four criteria and your decision: which review most closely resembles your question, which is most recent, which included most trials, and which is at lowest risk of bias. [4] Then check whether they disagree on direction or only on magnitude — in 86% of tested cases the direction agreed even when different reviews were chosen. [4] Disagreement on magnitude is usually explained by inclusion criteria and search dates.
Does this apply to overviews of reviews too?
Yes, and more sharply, because errors compound. The shockwave overview is an example: it appraised 8 reviews with AMSTAR 2 and their outcomes with GRADE, and concluded the treatment was effective while stating that the reliability of that conclusion is affected by the low methodological and evidential quality beneath it. [5] Reporting both halves of that sentence is the skill.
References
- Perry R, Whitmarsh A, Leach V, et al. A comparison of two assessment tools used in overviews of systematic reviews: ROBIS versus AMSTAR-2. Systematic Reviews. 2021 Oct 25;10(1):273. doi:10.1186/s13643-021-01819-x PMID 34696810 Inter-rater reliability study
- Buczinski S, Ferraro S, Vandeweerd JM. Assessment of systematic reviews and meta-analyses available for bovine and equine veterinarians and quality of abstract reporting: A scoping review. Preventive Veterinary Medicine. 2018 Dec 1;161:50–59. doi:10.1016/j.prevetmed.2018.10.011 PMID 30466658 Inter-rater reliability study
- Swierz MJ, Storman D, Zajac J, et al. Similarities, reliability and gaps in assessing the quality of conduct of systematic reviews using AMSTAR-2 and ROBIS: systematic survey of nutrition reviews. BMC Medical Research Methodology. 2021 Nov 27;21(1):261. doi:10.1186/s12874-021-01457-w PMID 34837960 Instrument comparison study
- Lunny C, Thirugnanasampanthar SS, Kanji S, et al. How can clinicians choose between conflicting and discordant systematic reviews? A replication study of the Jadad algorithm. BMC Medical Research Methodology. 2022 Oct 26;22(1):276. doi:10.1186/s12874-022-01750-2 PMID 36289496 Reproducibility study
- Zhou Q, Chen J. A critical overview of systematic reviews and meta-analyses of extracorporeal shockwave therapy for knee osteoarthritis. Asian Journal of Surgery. 2024 Jul;47(7):2975–2984. doi:10.1016/j.asjsur.2024.01.127 PMID 38290944 Overview of systematic reviews
About this resource
- Written by
- Dr Asif Khan (PT)BPT, MPT · Senior Physiotherapist · ShardaCare Healthcity, Greater Noida
- Reviewed by
- Dr Ravikant Mishra (PT)BPT, MPT · Head of Department · Department of Physiotherapy & Rehabilitation Science, ShardaCare Healthcity, Greater Noida · not the author
- Chief Editor
- Dr Dharam Pandey (PT)MPT; PhD
- Evidence grade
- ModerateSee "How certain is this?"
- Last reviewed
- 16 August 2026Next review due 16 August 2028
How to use this
Written to be learned from, not memorised.
This page separates three questions students routinely merge: how well the review was done, how certain its evidence is, and how large the effect was. Faculty may use this page in teaching with attribution. It carries its review date and its next review date, so you can see at a glance whether it is current before you put it in front of a cohort.
