Clinical library · Outcome measures
Modified Rankin Scale
Seven grades from no symptoms to death, and the primary endpoint of most modern stroke trials. Its central weakness is not what it measures but who is measuring: standard administration achieves an inter-rater kappa of 0.55.
In one line. A single ordinal grade describing global disability after stroke: 0 no symptoms, 1 symptoms without disability, 2 slight disability, 3 moderate disability requiring some help but able to walk unaided, 4 moderately severe, 5 severe, 6 dead.
It is the language stroke medicine argues in, and the reason to understand it well is that a one-grade disagreement between two clinicians is common enough to matter. The fix is known and rarely used.
The finding that should change practice. In a 2025 meta-analysis of 46 studies and 8,608 participants, overall inter-rater reliability was substantial (kappa 0.65, 95% CI 0.58 to 0.71) — but broken down, standard unstructured administration achieved only 0.55 (95% CI 0.46 to 0.64), while the Rankin Focussed Assessment reached 0.94 (95% CI 0.90 to 0.98). Introducing training raised reliability from 0.56 to 0.69. [1] If you use this scale untrained and unstructured, you are using its weakest form.
The numbers you actually need
| Property | Value | Source and caveat |
|---|---|---|
| Inter-rater reliability, overall | Kappa 0.65 (95% CI 0.58 to 0.71) | 29 studies within a review of 46; high risk of bias in 14 studies (30.4%) and possible publication bias [1] |
| Standard, unstructured | Kappa 0.55 (95% CI 0.46 to 0.64) | 13 studies [1]. An earlier synthesis reported the same pattern: 0.56 unstructured versus 0.78 with structured interviews [2] |
| Rankin Focussed Assessment | Kappa 0.94 (95% CI 0.90 to 0.98) | Only 2 studies — a large effect on a small base [1] |
| Test-retest reliability | Kappa 0.81 to 0.95 | Reported across the earlier literature synthesis of 50 selected articles [2] |
| Real-world agreement between wards | 70.5% agreement, kappa 0.55 | 105 patients transferred from stroke unit to rehabilitation ward. Certification of raters made little difference (kappa 0.57) [3] |
| Face-to-face versus telephone | Weighted kappa 0.80 (95% CI 0.74 to 0.87) | 5 studies [1]. Across 13 structured-questionnaire validation studies, weighted kappa ranged 0.56 to 0.90 but unweighted kappa only 0.27 to 0.68 [4] |
| Pre-stroke mRS | Weighted kappa 0.70 (95% CI 0.53 to 0.87) | 74 stroke survivors. Correlated well with a frailty index (rho 0.82) but poorly with medication count (0.28), and not at all with need for carers [5] |
What it measures
Global disability — the overall consequence of the stroke for the person's life, collapsed into one grade. That is its virtue as a trial endpoint: it captures everything at once, and it is meaningful to patients and clinicians without translation.
It is also why it is hard to score reliably. The judgement separating grade 2 from grade 3 is whether the person needs help with their usual activities, and that requires a consistent view of what counts as help and as usual. Structured instruments exist precisely to make that judgement uniform, and they work. [1][2]
Where it misleads
1. A kappa of 0.55 means routine one-grade disagreement
This is the practical fact. In real clinical practice, neurologists and rehabilitation physicians rating the same patients on the day of transfer agreed 70.5% of the time (kappa 0.55) — and being certified for trial use made almost no difference, at 70.0% and kappa 0.57. [3] The authors judge this sufficient for registries and observational studies, and their conclusion is a call for accessible training or predefined-question tools.
2. Structured and unstructured scores disagree most at the extremes
Across 13 validation studies of structured questionnaires, agreement with in-person unstructured assessment was generally good, but the discrepancies clustered at the ends of the scale: patients with good outcomes (mRS 2 or below) tended to rate themselves better than clinicians rated them, and those with poor outcomes (mRS 3 or above) tended to rate themselves worse. [4] Note also the gap between weighted kappa (0.56 to 0.90) and unweighted kappa (0.27 to 0.68) — weighting rewards near-misses, and near-misses are what this scale produces.
3. Pre-stroke mRS is used as a trial entry criterion and is only loosely valid
Interobserver reliability for pre-stroke mRS was 70% agreement, weighted kappa 0.70 — comparable with the standard scale. But validity was mixed: it correlated strongly with a frailty index (rho 0.82), moderately with comorbidity (0.50), weakly with medication count (0.28), and showed no association with need for carers. The authors conclude that relying on mRS alone may be a suboptimal measure of pre-stroke function and could bias trial samples. [5]
4. Utility-weighted versions can change a trial's answer
Utility weighting has been proposed to make the scale more patient-centred. A meta-analysis of 24 studies and 22,389 individuals found significant between-study variation in the weights assigned to each grade, with more variance at worse grades. When 18 major acute stroke trials were reanalysed using the different weighting scales, three produced an unstable outcome. [6] The weighting scheme is not a neutral technical choice.
What the evidence supports — and what it does not
Supported
- Validity as a global disability measure, with well-documented construct and convergent validity. [2]
- Structured assessment, training and adjudication to improve reliability. [1][2]
- Telephone administration, with substantial agreement against face-to-face. [1]
- Preference over the Barthel Index for detecting treatment effects in mild to moderate stroke. [7]
Not supported
- Unstructured administration as sufficient. Kappa 0.55. [1][3]
- Assuming certification alone fixes it. Certified raters agreed little better than uncertified ones. [3]
- Pre-stroke mRS as a stand-alone measure of prior function. [5]
- Treating any single utility-weighting scheme as neutral. Three of 18 reanalysed trials became unstable. [6]
- Reading unweighted agreement as good. It runs as low as 0.27. [4]
How certain is this?
Evidence grade: Moderate.
The reliability picture is the best-evidenced part and is consistent across two decades: a 2025 meta-analysis of 46 studies [1] reproduces the pattern described in a 2007 synthesis of 50 articles. [2] Independent replication over that interval is strong support for the central claim that structure and training matter more than anything else about this scale.
Two cautions on that meta-analysis, both stated by its authors: 30.4% of included studies were at high risk of bias, and publication bias is possible. [1] The headline figure for the Rankin Focussed Assessment rests on two studies. Its senior author also declares having contributed to commercially hosted mRS training materials and developed a remote mRS platform — a competing interest worth knowing when reading a paper that finds structured tools superior.
The real-world reliability study is a single centre with 105 assessed pairs. [3] The structured-questionnaire review found no validation data at all from Africa, and notes sampling bias and assessment delays as limitations. [4]
What would change the grade: multicentre reliability data for structured instruments, and validation outside high-income settings.
Common questions
Why do my colleague and I keep disagreeing by one grade?
Because that is what the scale does when administered conversationally. Standard unstructured administration pools to a kappa of 0.55, [1] and in a real stroke unit-to-rehabilitation transfer study two physicians agreed on 70.5% of patients. [3] The remedy is a structured instrument: the Rankin Focussed Assessment reached a kappa of 0.94 in the same meta-analysis, and simply introducing training raised reliability from 0.56 to 0.69. [1]
Does being certified for trials fix it?
Apparently not by itself. In the stroke unit study, agreement among certified physicians in both departments was 70.0% with a kappa of 0.57 — essentially the same as the overall figure. [3] The authors' conclusion is that easily accessible training or tools with predefined questions are what is needed.
Can I collect it by telephone?
Yes. Agreement between face-to-face and telephone administration was substantial, with a weighted kappa of 0.80 (95% CI 0.74 to 0.87) across five studies. [1] Be aware that across structured-questionnaire validation studies, discrepancies concentrated at the extremes — patients with good outcomes rated themselves better than clinicians did, and those with poor outcomes rated themselves worse. [4]
Should I record pre-stroke mRS?
Record it, but do not lean on it alone. Its interobserver reliability is comparable with the standard scale (weighted kappa 0.70), yet it correlated only weakly with medication count and not at all with need for carers, leading the authors to conclude that mRS alone may be a suboptimal measure of pre-stroke function that could bias trial samples. [5] It correlated strongly with frailty (rho 0.82), which is arguably what it is really capturing.
mRS or Barthel Index?
Different constructs: global disability against activities of daily living. For detecting treatment effects in mild to moderate stroke, a review of 44 studies found the mRS more sensitive and responsive, and noted that the Barthel's floor and ceiling effects may make it unsuitable for that purpose. [7] For describing how much help someone needs with self-care, the Barthel is the more direct measure — see the Barthel Index page.
References
- Tvrda L, Mavromati K, Taylor-Rowan M, et al. Comparing the properties of traditional and novel approaches to the modified Rankin scale: Systematic review and meta-analysis. European Stroke Journal. 2025 Jun;10(2):362–370. doi:10.1177/23969873241293569 PMID 39474736 Systematic review and meta-analysis
- Banks JL, Marotta CA. Outcomes validity and reliability of the modified Rankin scale: implications for stroke clinical trials: a literature review and synthesis. Stroke. 2007 Mar;38(3):1091–6. doi:10.1161/01.STR.0000258355.23810.c6 PMID 17272767 Literature review and synthesis
- Pożarowszczyk N, Kurkowska-Jastrzębska I, Sarzyńska-Długosz I, et al. Reliability of the modified Rankin Scale in clinical practice of stroke units and rehabilitation wards. Frontiers in Neurology. 2023;14:1064642. doi:10.3389/fneur.2023.1064642 PMID 36937517 Retrospective observational study
- Nair S, Hurly J, Saylor D. Evaluating the strengths and limitations of structured modified rankin scale validation studies - A systematic review. Journal of Stroke and Cerebrovascular Diseases. 2025 May;34(5):108242. doi:10.1016/j.jstrokecerebrovasdis.2025.108242 PMID 39922249 Systematic review
- Fearon P, McArthur KS, Garrity K, et al. Prestroke modified rankin stroke scale has moderate interobserver reliability and validity in an acute stroke setting. Stroke. 2012 Dec;43(12):3184–8. doi:10.1161/STROKEAHA.112.670422 PMID 23150650 Reliability and validity study
- Rebchuk AD, O'Neill ZR, Szefer EK, et al. Health Utility Weighting of the Modified Rankin Scale: A Systematic Review and Meta-analysis. JAMA Network Open. 2020 Apr 1;3(4):e203767. doi:10.1001/jamanetworkopen.2020.3767 PMID 32347948 Systematic review and meta-analysis
- Balu S. Differences in psychometric properties, cut-off scores, and outcomes between the Barthel Index and Modified Rankin Scale in pharmacotherapy-based stroke trials: systematic literature review. Current Medical Research and Opinion. 2009 Jun;25(6):1329–41. doi:10.1185/03007990902875877 PMID 19419341 Systematic literature review
About this resource
- Written by
- Dr Dharam Pandey (PT)MPT; PhD · Chief Editor · Director & Head of Department · Department of Physiotherapy & Rehabilitation Science
- Reviewed by
- Independent external peer reviewerAnonymous third-party review · not the author
- Evidence grade
- ModerateSee "How certain is this?"
- Last reviewed
- 16 August 2026Next review due 16 August 2028
Using this in clinic
Every figure here is traceable to its source.
Every kappa on this page is labelled with the method of administration that produced it, because for this scale the method is the main determinant of the number. Where a value could not be verified against the paper it came from, it is not on this page, and the omission is stated rather than filled with a number from a secondary source.
