Skip to content

Clinical library

Assessment and Outcome Measures

What each instrument measures, the change you can defend in a single patient, and where it misleads. Every reliability coefficient, cut-off and minimal clinically important difference is reported with the population it came from and the size of the study that produced it.

Published 16 measures· 89 verified sources

Why these pages read the way they do. Most outcome-measure references give you a number. The number is usually the least reliable part: an MCID derived from 35 people, a cut-off never validated in the population you work in, a measurement error quoted from a range whose lower half never applied to your patient. These pages give the number, and then the study behind it.

Where a value could not be verified against the paper that reported it, it is not here. That is why this section is short.

Balance, gait, function, disability and pain

Published now

Berg Balance Scale

Excellent rater reliability and a measurement error big enough to swallow most single-patient change. Explicitly not a falls screen on its own.

GRADE: Moderate7 sources

Timed Up and Go

A sound mobility measure and a poor falls screen. At the usual 13.5-second threshold, pooled sensitivity was 0.31.

GRADE: Moderate6 sources

Six-Minute Walk Test

The MCID depends entirely on the population. Published values run from 14 metres to 80, and both figures are correctly reported.

GRADE: Moderate6 sources

Functional Gait Assessment

Built to remove the Dynamic Gait Index ceiling, and it does. The famous 22/30 falls cut-off comes from 35 people.

GRADE: Low to moderate6 sources

Barthel Index

Excellent inter-rater reliability after stroke, and floor and ceiling effects the same literature says may stop it detecting treatment effects.

GRADE: Moderate7 sources

Modified Rankin Scale

The primary endpoint of most stroke trials. Administered conversationally it achieves an inter-rater kappa of 0.55.

GRADE: Moderate7 sources

Functional Independence Measure

Motor items are reliable, cognitive items much less so, and the cognitive subscore was not responsive to change where it was tested.

GRADE: Moderate5 sources

WOMAC

The standard measure in hip and knee osteoarthritis, with a pain subscale that shows no divergent validity from its own function subscale.

GRADE: Moderate7 sources

KOOS

Five subscales spanning acute injury to osteoarthritis. Choose the subscale before you collect it — relative efficiency differs threefold.

GRADE: Moderate5 sources

Oswestry Disability Index

Sixteen accepted methods gave MCIDs from 0.8 to 25 points in the same patients, and success rates from 30% to 83%.

GRADE: Moderate5 sources

Neck Disability Index

Good reliability across 79 studies. The important-change threshold has been reported anywhere from 5% to 33%.

GRADE: Moderate6 sources

DASH and QuickDASH

Pooled MCID about 11 points against a measurement error of about 9. The QuickDASH has strong negative evidence for responsiveness.

GRADE: Moderate5 sources

Patient-Specific Functional Scale

Reliable and highly responsive, and rated insufficient for construct validity as a measure of physical function.

GRADE: Moderate6 sources

Numeric Pain Rating Scale

About 2 points is the working threshold — but the change that counts is not uniform along the scale.

GRADE: Moderate to high6 sources

GMFM and GMFCS

One measures change and one explicitly does not. The COSMIN review says the GMFCS should not be used to detect change, and 56% kept the same rating throughout.

GRADE: Moderate5 sources

Expanded Disability Status Scale

Validity sufficient; reliability and sensitivity to change are documented weaknesses. Most rehabilitation trials sit in the band where the scale is driven by walking.

GRADE: Moderate5 sources

A note on citation age. Elsewhere on this site we prefer evidence published within about five years, because treatment evidence dates. The psychometric properties of an instrument do not date in the same way — the definitive reliability study for a scale published in 1989 may itself be fifteen years old and still be the best available. Where an older source is cited here, it is because it remains the primary derivation, and its age is visible in the reference list rather than hidden.

Using this in clinic

Every figure here is traceable to its source.

Each measure here is written the same way: what it tests, the numbers you would quote in a report, and the places where those numbers do not mean what they appear to. Where a value could not be verified against the paper it came from, it is not on this page, and the omission is stated rather than filled with a number from a secondary source.