Clinical library · Outcome measures
Barthel Index
Ten items of personal activities of daily living, scored 0 to 100. Excellent inter-rater reliability after stroke, and floor and ceiling effects severe enough that the same literature recommending it also warns it may not detect treatment effects.
In one line. Ten items — feeding, bathing, grooming, dressing, bowels, bladder, toilet use, transfers, mobility and stairs — scored by observation or report, giving a total that describes how much help a person needs with basic self-care.
It is quick, it is understood everywhere, and its weakness is at the two ends. A patient who is severely disabled and one who is profoundly so can both score near zero; a patient who is independent in self-care but cannot work, drive or manage a household scores 100.
"Barthel Index" is not one instrument. The clinimetric review is explicit that definitive statements about the scale are limited by heterogeneity in assessment methodology and in the content of the scales carrying the name, and it urges greater consistency in method, content and scoring. Its recommendation is that a 10-item scale scoring 0 to 100 in 5-point increments should become the uniform version for stroke trials. [2] Before comparing two Barthel scores, check they are the same Barthel.
The numbers you actually need
| Property | Value | Source and caveat |
|---|---|---|
| Inter-rater reliability, stroke | Weighted kappa 0.93 (95% CI 0.90 to 0.96) | 10 studies pooled, 543 participants. Individual studies ranged from 7 to 21 participants and only two were graded high quality [1] |
| Inter-observer disagreement, older people | Possibly around 4 points either way | 12 studies. No study investigated test-retest reliability at all, and item-level agreement was only fair to moderate [5] |
| Internal consistency | Cronbach alpha 0.96 (acute stroke trial data, n = 8,852) | The authors read this high value as evidence of redundancy in the scale, not as a virtue [6] |
| Reliability at ICU discharge | ICC 0.98 (95% CI 0.97 to 0.98); alpha 0.81 | 122 patients, two physiotherapists. Standard error of measurement 7.22 and smallest detectable change 20.01 points [4] |
| Responsiveness versus the modified Barthel | MBI better at group and individual level | Standardised response mean 1.10 versus 0.81 at first reassessment; the MBI detected significant improvement in 34.9% of patients against 25.4% for the standard BI [3] |
| Short form | 3 items: bladder control, transfer, mobility | Derived and externally validated; correlates at rho 0.83 or above with the full BI, modified Rankin Scale and NIHSS [6] |
| Versus the modified Rankin Scale | mRS more sensitive in mild to moderate stroke | 44 studies reviewed. The BI may not be appropriate for measuring treatment effects because of inherent floor and ceiling effects [7] |
The single most useful number here is the smallest detectable change of 20.01 points at ICU discharge. [4] On a 0-to-100 scale in 5-point steps, that is four whole increments — a reminder that an apparently large Barthel change in one patient can still be measurement noise.
What it measures
Basic personal activities of daily living, and nothing beyond them. There is no item for cooking, shopping, managing money, using transport, communication or cognition. That narrowness is deliberate and it is why the scale is fast; it is also why a score of 100 does not mean a person can live independently.
Scoring is by observation where possible, and by report from the patient or a carer where not. The reliability review in older people found evidence that the scale may be less reliable in patients with cognitive impairment, and when scores obtained by patient interview are compared with scores obtained by testing. [5] Those are exactly the circumstances in which it is most often collected by report.
Where it misleads
1. Floor and ceiling effects are the defining limitation
The clinimetric review states plainly that sensitivity to change is limited at the extremes of disability. [2] The consequence in trials is significant: a review of 44 studies concluded that the Barthel may not be an appropriate scale for measuring treatment effects because of its inherent ceiling and floor effects, and that in mild to moderate stroke the modified Rankin Scale appears to detect small treatment effects that the Barthel does not. [7]
Clinically, this is why a patient can plateau at 100 while still being unable to return to work or to a former role. The scale has stopped measuring, not the patient improving.
2. Excellent reliability was established on very small studies
The pooled weighted kappa of 0.93 is genuinely excellent, and the review that produced it notes that the individual studies ranged from 7 to 21 participants, that there was substantial clinical heterogeneity in how the scale was applied, and that only two papers were graded high quality. [1] The estimate is the best available; it is not a large evidence base.
3. Nobody has established its test-retest reliability in older people
The systematic review covering all the common clinical settings relevant to older people reports that no study investigated test-retest reliability, that inter-rater agreement for individual items was only fair to moderate, and that the role of assessor training on reliability has not been investigated. [5] Its conclusion is that important uncertainties remain. This is not the picture usually presented.
4. High internal consistency is a warning, not a recommendation
A Cronbach alpha of 0.96 across 8,852 acute stroke patients was interpreted by the analysts as evidence of redundancy: several items are carrying the same information. They derived and validated a three-item version — bladder control, transfer and mobility — that correlates with the full scale at rho of 0.83 or above. [6] If three items reproduce ten, the other seven are not adding measurement.
What the evidence supports — and what it does not
Supported
- Excellent inter-rater reliability after stroke for standard administration. [1]
- Use as a stroke trial outcome, with the 10-item 0-to-100 version as the uniform standard. [2]
- Reliable and valid assessment at ICU discharge, with a smallest detectable change of 20.01 points. [4]
- Preferring the modified Barthel where detecting change matters. [3]
- A validated three-item short form where brevity is needed. [6]
Not supported
- Detecting change at the extremes of disability. [2][7]
- Using it as the primary measure of treatment effect in mild to moderate stroke. The mRS is more sensitive. [7]
- Assuming test-retest reliability in older people. It has never been studied. [5]
- Comparing scores across differently-constructed "Barthel" scales. [2]
- Reading a score of 100 as independent living. The scale contains no instrumental activity.
How certain is this?
Evidence grade: Moderate.
The reliability finding after stroke is a proper meta-analysis with a tight confidence interval, [1] and it is corroborated independently in intensive care. [4] Those are the firmest claims here.
Everything about responsiveness is weaker. The comparison against the modified Barthel comes from a single centre with 63, 63 and 55 patients at three time points. [3] The judgement that the mRS outperforms the Barthel for treatment effects comes from a narrative systematic review that also notes inconsistency in cut-off points across studies and calls for more research to differentiate the two scales. [7]
The gaps identified in older people are gaps in the literature rather than demonstrated failures, and the review says so. [5] The short-form work is methodologically strong, drawing on 8,852 acute and 332 rehabilitation trial participants. [6]
What would change the grade: test-retest data in older people, and MCID values derived in populations other than critical care.
Common questions
What change can I defend in one patient?
The best-quantified figure is a smallest detectable change of 20.01 points, with a standard error of measurement of 7.22, derived at intensive care discharge in 122 patients. [4] On a scale that moves in 5-point steps, that is four increments. Outside critical care, published minimal detectable change values for the standard Barthel are not well established, and this page will not supply one it cannot source.
Barthel or modified Barthel?
The modified version, if you need to detect change. In early subacute stroke it showed significantly larger standardised response means (1.10 against 0.81) and identified significantly more patients with meaningful improvement (34.9% against 25.4%). The authors recommend it for clinical and research use, while noting the two are equally responsive when the improvement is large. [3]
Barthel or modified Rankin Scale for a stroke trial?
They measure different things — the Barthel measures activities of daily living, the mRS global disability. A review of 44 studies concluded the mRS appears more sensitive and responsive for stroke disability, particularly in mild to moderate stroke, and that the Barthel's floor and ceiling effects may make it unsuitable for detecting treatment effects. [7] See the section index for the mRS page.
Can I use a shortened version?
Yes, and there is a validated one. Factor analysis of acute stroke trial data identified an optimal three-item short form comprising bladder control, transfer and mobility, which correlated with the full Barthel, the mRS and the NIHSS at rho of 0.83 or above. [6] The same analysis found the full scale internally redundant, with an alpha of 0.96.
Is it reliable if a relative reports the answers?
Less so, and this is under-studied. The review in older people found evidence that the scale may be less reliable in patients with cognitive impairment, and when interview scores are compared with scores from actual testing. It also found no study of test-retest reliability at all. [5] Record how you collected it.
References
- Duffy L, Gajree S, Langhorne P, et al. Reliability (inter-rater agreement) of the Barthel Index for assessment of stroke survivors: systematic review and meta-analysis. Stroke. 2013 Feb;44(2):462–8. doi:10.1161/STROKEAHA.112.678615 PMID 23299497 Systematic review and meta-analysis
- Quinn TJ, Langhorne P, Stott DJ. Barthel index for stroke trials: development, properties, and application. Stroke. 2011 Apr;42(4):1146–51. doi:10.1161/STROKEAHA.110.598540 PMID 21372310 Narrative review
- Wang YC, Chang PF, Chen YM, et al. Comparison of responsiveness of the Barthel Index and modified Barthel Index in patients with stroke. Disability and Rehabilitation. 2023 Mar;45(6):1097–1102. doi:10.1080/09638288.2022.2055166 PMID 35357990 Comparative responsiveness study
- Dos Reis NF, Figueiredo FCXS, Biscaro RRM, et al. Psychometric Properties of the Barthel Index Used at Intensive Care Unit Discharge. American Journal of Critical Care. 2022 Jan 1;31(1):65–72. doi:10.4037/ajcc2022732 PMID 34972844 Observational psychometric study
- Sainsbury A, Seebass G, Bansal A, et al. Reliability of the Barthel Index when used with older people. Age and Ageing. 2005 May;34(3):228–32. doi:10.1093/ageing/afi063 PMID 15863408 Systematic review
- MacIsaac RL, Ali M, Taylor-Rowan M, et al. Use of a 3-Item Short-Form Version of the Barthel Index for Use in Stroke: Systematic Review and External Validation. Stroke. 2017 Mar;48(3):618–623. doi:10.1161/STROKEAHA.116.014789 PMID 28154094 Systematic review and external validation
- Balu S. Differences in psychometric properties, cut-off scores, and outcomes between the Barthel Index and Modified Rankin Scale in pharmacotherapy-based stroke trials: systematic literature review. Current Medical Research and Opinion. 2009 Jun;25(6):1329–41. doi:10.1185/03007990902875877 PMID 19419341 Systematic literature review
About this resource
- Written by
- Dr Dharam Pandey (PT)MPT; PhD · Chief Editor · Director & Head of Department · Department of Physiotherapy & Rehabilitation Science
- Reviewed by
- Independent external peer reviewerAnonymous third-party review · not the author
- Evidence grade
- ModerateSee "How certain is this?"
- Last reviewed
- 16 August 2026Next review due 16 August 2028
Using this in clinic
Every figure here is traceable to its source.
Every reliability coefficient and change threshold on this page carries the population and the sample size behind it, because for this scale the version used and the setting change the answer. Where a value could not be verified against the paper it came from, it is not on this page, and the omission is stated rather than filled with a number from a secondary source.
