Skip to content

Clinical library · Outcome measures

Berg Balance Scale

The most widely used clinical balance scale, with excellent rater reliability and a measurement error large enough to swallow most single-patient improvements. Both of those things are true at once, and the second is the one that changes practice.

Evidence 14 items · 0–56 points· Reliability meta-analysis, 65 studies· MCID varies by population

In one line. Fourteen functional tasks, each scored 0 to 4, giving a total out of 56. It measures static and anticipatory balance during everyday tasks, it takes about 15 to 20 minutes, and it needs a chair, a step and a ruler.

The scale's problem is not its reliability, which is excellent. It is the gap between that reliability and the size of change you can defend in an individual patient — and the ceiling effect that hides improvement in the people who are doing well.

The numbers you actually need

PropertyValueSource and caveat
Intra-rater reliability Pooled 0.98 (95% CI 0.97 to 0.99) 11 studies, 668 participants [1]. A separate reliability generalization meta-analysis of 80 samples in 65 studies gave a mean intra-rater ICC of 0.957 and inter-rater 0.97 [2]
Inter-rater reliability Pooled 0.97 (95% CI 0.96 to 0.98) Same review [1]
Internal consistency Mean alpha 0.92 (range 0.62 to 0.98) Across 80 samples; reliability is a property of the application, not of the instrument, so it should be reported per study [2]
Minimal detectable change (95%) 2.8 to 6.6 points Varies across the scale, and all contributing studies had mean scores of 20 or above out of 56 [1]. A single-centre study in mixed balance disorders gave SEM 2.18 and MDC95 6.2 points [3]
MCID, multiple sclerosis 3 points 110 people, anchored to a 10% improvement on the Activities-specific Balance Confidence Scale; area under the curve 0.65 [4]
MCID, early subacute stroke 5 points if walking assistance is needed 80 patients; AUC 0.84, p < .01. In the unassisted group the estimate was 4 points with an AUC of 0.62 and p = .26 — not a usable threshold [5]
Minimal important change, mixed balance disorders 7 points Anchor-based, against a 15-point Global Rating of Change [3]
Healthy older adults Decline of 0.7 points per year 17 studies, 1,363 participants; mean scores 37 to 55, and variability rises with age [7]

Note the arithmetic that follows from rows four to seven: in multiple sclerosis the published MCID of 3 points [4] sits below the lower bound of the measurement error reported for the scale [1]. That is not a contradiction in the literature so much as a warning about what a 3-point change in one patient can carry.

What it measures

Fourteen tasks, ordered roughly by difficulty: sitting to standing, standing unsupported, sitting unsupported, standing to sitting, transfers, standing with eyes closed, standing with feet together, reaching forward with an outstretched arm, retrieving an object from the floor, turning to look behind, turning 360 degrees, placing alternate feet on a step, standing with one foot in front, and standing on one leg. Each is scored 0 to 4 against written criteria.

What that set covers is static and anticipatory postural control in functional positions. What it does not contain is any reactive task — nothing in the scale perturbs the patient and asks them to recover. It is a measure of balance maintained, not balance regained.

Where it misleads

1. The ceiling effect is not a footnote

A ceiling effect was evident for some participants in the reliability review, and the scale has higher absolute reliability close to 56 points precisely because of it. [1] In a head-to-head comparison after 10 sessions of physical therapy, 12 of 93 participants (12.9%) hit the maximum score on the Berg, against 2 (2.1%) on the Mini-BESTest. [3]

The practical consequence: a patient scoring in the high 40s or 50s who is still falling is not well served by this scale. Their improvement has nowhere to go.

2. It was never validated at the bottom of its own range

This is the finding least often quoted. Every study contributing to the absolute reliability analysis had an average score of 20 or above out of 56, and the reviewers state explicitly that they identified no data estimating absolute reliability among participants with a mean score below 20. [1] If your patient scores 12, the published measurement error does not apply to them, because nobody has measured it there.

3. It is a poor falls predictor on its own

A systematic review of eight prognostic studies found mean Berg scores were high regardless of falls history, only three studies offered cut-off scores at all, and those ranged from 45 to 51 points. Its conclusion is unambiguous: the evidence supporting the use of the Berg to predict falls is insufficient, and it should not be used alone to determine falls risk in older adults. [6]

The scale is routinely used exactly that way. If falls risk is the question, it belongs in a multifactorial assessment — see falls prevention in older adults, where the intervention evidence is much stronger than the screening evidence.

4. "Reliable" describes the study, not your clinic

The reliability generalization meta-analysis makes a methodological point with direct practical force: reliability is not a fixed property of a scale but of the results obtained with it, and it varies with the context and the participants. Alpha ranged from 0.62 to 0.98 across 80 samples. [2] Quoting "ICC 0.98" as though it were a specification of the instrument is a misreading.

What the evidence supports — and what it does not

Supported

  • Excellent rater reliability for group and repeated measurement. [1][2]
  • Population-specific MCIDs where they have been derived — 3 points in multiple sclerosis, [4] 5 points in assisted-walking early subacute stroke. [5]
  • Tracking change over a course of rehabilitation in patients scoring in the middle of the range. [1][3]
  • Normative comparison in healthy older adults, with the expectation of decline after 70. [7]

Not supported

  • Predicting falls on its own. Explicitly recommended against. [6]
  • Applying the published measurement error below a score of 20. No data exist there. [1]
  • Detecting improvement near the ceiling. [1][3]
  • Treating a single MCID as universal. Values differ by population, and one was derived with an AUC of 0.62 and p = .26. [4][5]
  • Assessing reactive balance. No item tests it.

How certain is this?

Evidence grade: Moderate.

The reliability estimates are the most secure figures on this page: two independent meta-analyses, 11 and 65 studies respectively, agreeing closely. [1][2]

The MCID estimates are weaker than their widespread quotation suggests. The multiple sclerosis value rests on 110 people with an area under the curve of 0.65 [4] — modest discrimination. The stroke value is stronger in the assisted-walking group (AUC 0.84) and should not be used in the unassisted group, where the same study reported an AUC of 0.62 and a non-significant p value of .26. [5] Both are single studies in single health systems.

The falls-prediction conclusion is a narrative synthesis of eight studies that could not be pooled because of heterogeneity, with low to moderate risk of bias. [6] It is a statement that the evidence is insufficient, not a demonstration that the scale fails. The distinction matters and this site keeps it.

What would change the grade: MDC data from patients scoring below 20, and MCID replication outside single centres.

Common questions

What change can I defend in one patient?

Depends where they started and which population they belong to. The pooled minimal detectable change with 95% confidence ranges from 2.8 to 6.6 points and varies across the scale. [1] A single-centre study in mixed balance disorders reported MDC 6.2 points and a minimal important change of 7. [3] If you need one working rule: a change smaller than about 6 points in an individual is inside the measurement error reported by most studies.

Should I use the Berg or the Mini-BESTest?

In the direct comparison of 93 participants, the two behaved similarly, but the Mini-BESTest had a lower ceiling effect, slightly higher test-retest reliability (ICC 0.96 against 0.92) and greater accuracy in classifying individual patients who improved. After treatment, 38 participants assessed with the Mini-BESTest showed a change at or above the minimal important change, against 23 assessed with the Berg. [3] For higher-functioning patients that difference is the whole argument.

Can I use it to decide who needs a falls programme?

Not on its own. A systematic review found the evidence insufficient and stated the scale should not be used alone to determine falls risk in older adults; the three cut-offs that exist range from 45 to 51 points. [6] Use it as one component of a multifactorial assessment.

What is a normal score for an older adult?

Higher than most people assume, and falling with age. Across 17 studies and 1,363 healthy community-dwelling people with a mean age of 70 or more, mean scores ranged from 37 to 55 out of 56, with a decline of 0.7 points per year and rising variability with age. [7] A score in the low 50s in an 85-year-old is not necessarily abnormal.

My patient scores 8. What is their measurement error?

Unknown. Every study in the absolute reliability analysis had a mean score of 20 or above, and the reviewers found no data below that. [1] Report the score, do not attach a published MDC to it, and consider whether a scale with items at that difficulty level would serve better.

References

  1. Downs S, Marquez J, Chiarelli P. The Berg Balance Scale has high intra- and inter-rater reliability but absolute reliability varies across the scale: a systematic review. Journal of Physiotherapy. 2013 Jun;59(2):93–9. doi:10.1016/S1836-9553(13)70161-9 PMID 23663794 Systematic review with meta-analysis
  2. Meseguer-Henarejos AB, Rubio-Aparicio M, López-Pina JA, et al. Characteristics that affect score reliability in the Berg Balance Scale: a meta-analytic reliability generalization study. European Journal of Physical and Rehabilitation Medicine. 2019 Oct;55(5):570–584. doi:10.23736/S1973-9087.19.05363-2 PMID 30955319 Reliability generalization meta-analysis
  3. Godi M, Franchignoni F, Caligari M, et al. Comparison of reliability, validity, and responsiveness of the mini-BESTest and Berg Balance Scale in patients with balance disorders. Physical Therapy. 2013 Feb;93(2):158–67. doi:10.2522/ptj.20120171 PMID 23023812 Prospective observational study
  4. Gervasoni E, Jonsdottir J, Montesano A, et al. Minimal Clinically Important Difference of Berg Balance Scale in People With Multiple Sclerosis. Archives of Physical Medicine and Rehabilitation. 2017 Feb;98(2):337–340.e2. doi:10.1016/j.apmr.2016.09.128 PMID 27789239 Cohort study
  5. Tamura S, Miyata K, Kobayashi S, et al. The minimal clinically important difference in Berg Balance Scale scores among patients with early subacute stroke: a multicenter, retrospective, observational study. Topics in Stroke Rehabilitation. 2022 Sep;29(6):423–429. doi:10.1080/10749357.2021.1943800 PMID 34169808 Multicentre retrospective observational study
  6. Lima CA, Ricci NA, Nogueira EC, et al. The Berg Balance Scale as a clinical screening tool to predict fall risk in older adults: a systematic review. Physiotherapy. 2018 Dec;104(4):383–394. doi:10.1016/j.physio.2018.02.002 PMID 29945726 Systematic review
  7. Downs S, Marquez J, Chiarelli P. Normative scores on the Berg Balance Scale decline after age 70 years in healthy community-dwelling people: a systematic review. Journal of Physiotherapy. 2014 Jun;60(2):85–9. doi:10.1016/j.jphys.2014.01.002 PMID 24952835 Systematic review with meta-analysis

About this resource

Using this in clinic

Every figure here is traceable to its source.

Every reliability coefficient, measurement error and MCID on this page was read off the paper that reported it, with the population and the method of derivation stated alongside it. Where a value could not be verified against the paper it came from, it is not on this page, and the omission is stated rather than filled with a number from a secondary source.