Clinical library · Outcome measures
Patient-Specific Functional Scale
The patient names the activities that matter to them and rates each out of ten. Reliable and highly responsive across many conditions — and, on COSMIN appraisal, of insufficient construct validity as a measure of physical function.
In one line. The patient identifies up to five activities they are unable to do or have difficulty with because of their problem, and rates current ability on each from 0 (unable) to 10 (able at pre-injury level). The score is the average.
Its appeal is obvious: it measures what the patient actually came in about, and it takes two minutes. The complication is that an instrument whose content changes with every patient is difficult to validate as a measure of anything in particular.
Reliable and responsive; construct validity uncertain. The COSMIN review of 57 measurement-property studies found sufficient test-retest reliability in musculoskeletal conditions (22 studies, 845 participants, low-to-moderate certainty) and sufficient responsiveness (32 studies, 13,770 participants, moderate-to-high certainty) — but insufficient construct validity as a measure of physical function (21 studies, 2,945 participants, low-to-moderate certainty). It also found the scale in use across 87 health conditions, some without prior evidence of validity. [1]
The numbers you actually need
| Property | Value | Source and caveat |
|---|---|---|
| Standard error of measurement | 0.35 to 1.5 points | Across the 57 reviewed studies [1] |
| Neck pain (PSFS 2.0) | ICC 0.95; smallest detectable change 1.10; minimal important change 2.67 | 100 patients; no floor or ceiling effect; AUC 0.82 [2] |
| Subacute stroke rehabilitation | ICC 0.81; SEM 0.70; smallest detectable change 1.94; minimal important change 1.58 | Note the minimal important change is below the smallest detectable change. Ceiling effect in 25% of participants at 3 months post-discharge [4] |
| Multiple sclerosis, mobility goals | ICC 0.70; MDC 2.1; MCID 2.5 or more | Sensitivity 0.85, specificity 0.76 against a global rating of change; responsive (d = 1.7) [5] |
| Total knee arthroplasty | ICC 0.73 to 0.86; effect size 1.71 at 3 months, 2.89 at 1 year | Larger effect sizes than the WOMAC at both points, but limits of agreement ±2.17 to ±2.72 [6] |
| Upper extremity disorders | ICC 0.75 or above in shoulder pain, multiple shoulder disorders and hand osteoarthritis | Construct validity only moderate: r = 0.50 against the Upper Extremity Functional Index and 0.51 against the NPRS [3] |
What it measures
Whatever the patient says matters. That is the design, and it is why the scale correlates only moderately with fixed-item functional measures — it is not trying to measure the same thing they are.
The practical consequence is that the score is only interpretable alongside the activities chosen. A change from 3 to 7 on "carry my grandchild" and the same change on "sleep through the night" are not the same clinical event, and the number alone does not distinguish them. Record the activities, not just the score.
Where it misleads
1. Insufficient construct validity is a real finding, not a technicality
Across 21 studies and 2,945 participants, the COSMIN appraisal rated construct validity as a measure of physical function insufficient. [1] In upper extremity disorders, correlations were only moderate — 0.50 with a fixed-item functional index and 0.51 with a pain rating. [3] If you need a measure of physical function that will stand up as such, this is not it. If you need a measure of progress towards the patient's own goals, it is excellent.
2. In stroke, the important change is smaller than the detectable change
The subacute stroke study reported a minimal important change of 1.58 points against a smallest detectable change of 1.94. [4] The threshold patients regard as meaningful sits inside the measurement error of the instrument, which means an individual change at that level cannot be distinguished from noise. Both numbers are correctly derived; together they define a limit on individual-level interpretation.
3. Thresholds differ by population by a factor of nearly two
Minimal important change was 2.67 in neck pain, [2] 2.5 or more in multiple sclerosis, [5] and 1.58 in subacute stroke. [4] These are not interchangeable, and the two-point rule of thumb often quoted sits awkwardly across all three.
4. It has spread far ahead of its validation
The review found the scale used in 87 unique health conditions, some without prior evidence of validity, and concluded that further study in non-musculoskeletal conditions is necessary before clinical use. [1] Reliability evidence outside musculoskeletal practice rested on 6 studies and 197 participants at very low certainty.
5. Ceiling effects appear once patients recover
A quarter of stroke participants showed a ceiling effect three months after discharge. [4] Because the patient nominates activities at their worst, those activities become easy as they improve — and the scale stops discriminating unless the activities are renegotiated.
What the evidence supports — and what it does not
Supported
- Sufficient test-retest reliability in musculoskeletal conditions. [1][3]
- Sufficient responsiveness, on moderate-to-high certainty across 32 studies and 13,770 participants. [1]
- Tracking progress towards patient-nominated goals, including after stroke and in multiple sclerosis. [4][5]
- Larger effect sizes than the WOMAC after knee arthroplasty. [6]
Not supported
- Use as a validated measure of physical function. Construct validity insufficient. [1]
- Use in non-musculoskeletal conditions without further study. The review says so explicitly. [1]
- A single two-point threshold across populations. Values run 1.58 to 2.67. [2][4][5]
- Interpreting individual change in stroke at the published minimal important change. It is below the measurement error. [4]
- Using it interchangeably with the WOMAC early after arthroplasty. [6]
How certain is this?
Evidence grade: Moderate.
The evidence base is unusually large for a scale this simple — 57 measurement-property studies and 255 papers describing its use, appraised with COSMIN and GRADE by two independent reviewers, preregistered. [1] Its findings are graded rather than asserted, which is why this page reports "sufficient", "insufficient" and their certainty levels rather than a summary verdict.
The individual population studies are single-centre cohorts of 100 patients or fewer, [2][4][5][6] which is why the thresholds vary. The upper-limb review covers 14 studies of which three were of adequate to very good quality. [3]
What would change the grade: an agreed answer on what construct the scale measures, and minimal important change values derived consistently across populations rather than one cohort at a time.
Common questions
Is a two-point change meaningful?
Usually, but check your population. Minimal important change was 2.67 in neck pain, [2] 2.5 or more in multiple sclerosis, [5] and 1.58 in subacute stroke — where it sits below the smallest detectable change of 1.94, so an individual change at that level cannot be separated from measurement error. [4] The standard error of measurement across the literature ranges from 0.35 to 1.5 points. [1]
Can I use it as my measure of physical function?
Not if it has to stand up as one. The COSMIN review rated construct validity as a measure of physical function insufficient across 21 studies and 2,945 participants. [1] Correlations with fixed-item functional measures are moderate at best — 0.50 against the Upper Extremity Functional Index. [3] Use it for goal attainment and pair it with a condition-specific measure if you need function.
Does it work in neurological conditions?
The specific evidence is encouraging and the general advice is cautious. In subacute stroke it showed satisfactory content validity, ICC 0.81 and high responsiveness; [4] in multiple sclerosis, moderate reliability (ICC 0.70), responsiveness of d = 1.7 and an MCID of 2.5 or more. [5] But the COSMIN review found reliability in non-musculoskeletal conditions rested on 6 studies and 197 participants at very low certainty, and recommends further study before clinical use. [1]
My patient's score stopped moving. Have they plateaued?
Possibly they have outgrown their own activities. A ceiling effect was found in 25% of stroke participants three months after discharge. [4] Because activities are nominated at the point of greatest difficulty, they become easy with recovery. Renegotiate the list rather than concluding progress has stopped.
Should I record the activities or just the score?
The activities. The number is only interpretable with them, and the review found the scale used across 87 conditions with varying content, some without prior evidence of validity. [1] The activity list is the audit trail for what the score actually refers to.
References
- Pathak A, Wilson R, Sharma S, et al. Measurement Properties of the Patient-Specific Functional Scale and Its Current Uses: An Updated Systematic Review of 57 Studies Using COSMIN Guidelines. Journal of Orthopaedic & Sports Physical Therapy. 2022 May;52(5):262–275. doi:10.2519/jospt.2022.10727 PMID 35128944 Systematic review (COSMIN)
- Thoomes E, Cleland JA, Falla D, et al. Reliability, Measurement Error, Responsiveness, and Minimal Important Change of the Patient-Specific Functional Scale 2.0 for Patients With Nonspecific Neck Pain. Physical Therapy. 2024 Jan 1;104(1):. doi:10.1093/ptj/pzad113 PMID 37606246 Prospective cohort study
- Nazari G, Bobos P, Lu Z, et al. Psychometric properties of Patient-Specific Functional Scale in patients with upper extremity disorders. A systematic review. Disability and Rehabilitation. 2022 Jun;44(13):2958–2967. doi:10.1080/09638288.2020.1851784 PMID 33290102 Systematic review
- Evensen J, Soberg HL, Sveen U, et al. Measurement Properties of the Patient-Specific Functional Scale in Rehabilitation for Patients With Stroke: A Prospective Observational Study. Physical Therapy. 2023 May 4;103(5):. doi:10.1093/ptj/pzad014 PMID 37140476 Prospective observational study
- Mañago MM, Cohen ET, Cameron MH, et al. Reliability, Validity, and Responsiveness of the Patient-Specific Functional Scale for Measuring Mobility-Related Goals in People With Multiple Sclerosis. Journal of Neurologic Physical Therapy. 2023 Jul 1;47(3):139–145. doi:10.1097/NPT.0000000000000439 PMID 36897202 Prospective cohort study
- Berghmans DD, Lenssen AF, van Rhijn LW, et al. The Patient-Specific Functional Scale: Its Reliability and Responsiveness in Patients Undergoing a Total Knee Arthroplasty. Journal of Orthopaedic & Sports Physical Therapy. 2015 Jul;45(7):550–6. doi:10.2519/jospt.2015.5825 PMID 25996364 Prospective cohort study
About this resource
- Written by
- Dr Anuj Mishra (PT)BPT, MPT · Head of Department · Department of Physiotherapy & Rehabilitation Science, Shanti Mukand Hospital, Karkardooma, Delhi
- Reviewed by
- Dr Dharam Pandey (PT)MPT; PhD · Chief Editor · Director & Head of Department · Department of Physiotherapy & Rehabilitation Science · not the author
- Chief Editor
- Dr Dharam Pandey (PT)MPT; PhD
- Evidence grade
- ModerateSee "How certain is this?"
- Last reviewed
- 16 August 2026Next review due 16 August 2028
Using this in clinic
Every figure here is traceable to its source.
Every threshold on this page names its population, because for a scale whose content changes with every patient, the population is most of the meaning. Where a value could not be verified against the paper it came from, it is not on this page, and the omission is stated rather than filled with a number from a secondary source.
