Student library · Practical skills
Goniometry
On a printed protractor the goniometer agrees with itself to within about a degree. On a human shoulder the 95% limits of agreement widen to roughly ±11°. You will be taught to record range of motion to the nearest degree; almost none of those degrees are real.
In one line. A goniometer measures joint angle. Its reliability is usually excellent as a correlation and considerably worse as an error in degrees, and those two facts are reported in the same papers.
The skill you are being assessed on is placement and stabilisation. The skill you will need in clinic is knowing how much of a recorded change is your patient and how much is your hands.
A high ICC is not a small error. The same study reports intra-rater ICC up to 0.996 for human joint angles and a minimal detectable change of up to 9° for the same measurements. [1] An ICC tells you how well the measure separates people who differ; the standard error of measurement and the minimal detectable change tell you what a change in one patient has to exceed before you can call it real. Quote the second pair.
The numbers you actually need
| Question | Finding | Source and caveat |
|---|---|---|
| How good is it on a known angle? | Excellent. Across five devices, ICC above 0.98, standard error of measurement 0.59° to 1.75°, minimal detectable change 1° to 4° | 12 standard angles, 3 examiners, 2 sessions. 95% limits of agreement -4.11° to 4.04° [1] |
| How good is it on a person? | Much worse. Inter-rater ICC 0.697 to 0.975, SEM 1.93° to 4.64°, minimal detectable change 5° to 11°. Intra-rater ICC 0.660 to 0.996, SEM 0.77° to 4.06°, MDC 2° to 9° | Same study, 160 shoulder measurements from 20 shoulders. 95% limits of agreement widened to -10.98° to 11.36° [1] |
| Does it get worse at end range? | Yes. Wider angle measurement resulted in reduced device reliability | Measured across 0° to 180°. The authors recommend evaluating error separately across the range rather than quoting one figure [1] |
| What about the knee? | Intra-rater ICC 0.997 to 0.998 with smallest detectable difference 3.19° to 4.09°; inter-rater ICC 0.994 with SDD 5.85° | Digital goniometer. The authors note the relatively high SDD "may indicate a problem monitoring small differences" — alongside an ICC of 0.994 [2] |
| Does experience help? | Not detectably. No statistical difference between a novice and an experienced assessor (p = 0.86) | Knee ROM, digital goniometer [2]. In the elbow review, one study that gave instructions on alignment found no expert/non-expert difference, while one that gave none found experts more reliable [3] |
| How wide is the reliability range in the literature? | Intra-rater ICC 0.45 to 0.99; inter-rater 0.53 to 0.97 | 12 studies of the elbow, 6 rated high quality. One validity study found a maximum standard error of the mean of 11.5° for total range of motion [3] |
| Do we know the measurement error properly? | Often not. Of 15 studies of wrist goniometry, 1 was of fair methodological quality and 14 were poor; the level of evidence for measurement error was unknown because minimal important change was never calculated | COSMIN appraisal. "Limited" evidence for acceptable reliability even in the best method [4] |
| Does technique change the number? | Yes. Reliability was good to excellent for shoulder rotation (ICC 0.85 to 0.99), but systematic differences appeared across trials and between testers, and patient position and equipment produced different outcome measures | Which is why a department protocol matters more than an individual's skill [5] |
Where students get this wrong
1. Recording to the degree and believing the degree
If the minimal detectable change for a human joint angle runs from 5° to 11° between raters, [1] then "flexion improved from 112° to 118°" is not a finding. It is inside the noise. Record what you measured, but interpret change against the error, and say so in your notes.
2. Quoting the ICC because it is the flattering number
The knee study reports inter-rater ICC of 0.994 and a smallest detectable difference of 5.85° in the same result. [2] Both are correct. The ICC is high because the sample spans a wide range of knee angles, which is what an ICC rewards; the SDD is what applies to your one patient. The authors themselves flag the tension. An exam answer that gives the ICC alone is incomplete.
3. Assuming your measurements will improve with practice
They will improve in placement and speed. Whether they improve in reliability is less clear: a novice and an experienced assessor did not differ detectably (p = 0.86), [2] and in the elbow review the expert advantage appeared only in the study where examiners were given no instructions on alignment. [3] The implication is uncomfortable and useful — a written protocol closes the gap that experience does not.
4. Comparing your number to someone else's
Inter-rater error is consistently worse than intra-rater error: MDC 5° to 11° versus 2° to 9° in the same study. [1] If a colleague measured at admission and you measure at discharge, the threshold for real change is the larger one. Where it matters, re-measure yourself rather than trusting the earlier figure.
5. Trusting the equipment comparison you read online
Across five devices, concurrent validity was excellent (ICC above 0.99) — on standard angles. On human joints the limits of agreement widened to nearly ±11°. [1] The device is rarely the limiting factor; the joint, the soft tissue and the end-feel are. The study's own recommendation is to assess measurement error separately for standard and human angles, which is precisely the distinction most equipment reviews collapse.
6. Assuming the literature is solid because the tool is old
In wrist goniometry, 14 of 15 studies were of poor methodological quality and the evidence for measurement error was rated unknown. [4] A technique being a century old and taught everywhere is not the same as its error being characterised. This is the same lesson as the red flags literature, in a different corner of practice.
What the evidence supports — and what it does not
Supported
- Goniometry as a reliable clinical measure when error is interpreted properly. [1][2][5]
- The same rater measuring at both time points. Intra-rater error is smaller. [1]
- A written departmental protocol for position and landmarks. [3][5]
- Reporting SEM or MDC alongside any ICC. [1][2]
Not supported
- Treating a change of a few degrees as clinical improvement. [1][2]
- Quoting standard-angle accuracy as if it applied to patients. [1]
- Assuming experience alone improves reliability. [2][3]
- A single error figure across the whole range. Reliability fell at wider angles. [1]
- Assuming wrist goniometry error is known. It is rated unknown. [4]
How certain is this?
Evidence grade: Moderate.
The core finding — that error on human joints is several times larger than error on known angles — is measured directly in a study designed to separate the two, across five devices, three examiners and two sessions. [1] It is consistent with the knee and elbow data. [2][3]
What limits the grade is sample size and quality. The device study measured 20 shoulders; the knee study is single-centre; the elbow review found only 6 of 12 studies high quality, [3] and the wrist review found 14 of 15 poor. [4] The direction of the finding is secure; the precise thresholds for any given joint are not.
What would change the grade: adequately powered reliability studies per joint that report minimal detectable change and minimal important change together, so that "measurable" and "meaningful" can finally be separated.
Common questions
How much change in range of motion is real?
Depends on the joint, the rater and where in the range you are. As a working figure from the device study: 2° to 9° if you measure the patient both times, and 5° to 11° if someone else took the first measurement. [1] For the knee with a digital goniometer, roughly 3° to 4° within a rater and 5.85° between raters. [2] Anything smaller cannot be distinguished from measurement error.
Is a smartphone app acceptable?
On this evidence, yes — the smartphone application showed superior reliability for human joint angles among the five devices tested, while the digital inclinometer was best on standard angles. [1] Concurrent validity between all device pairs was excellent. The device matters less than the position, the landmarks and whether the same person measures each time. [5]
Do I need to worry about this in my OSCE?
Your OSCE marks placement, stabilisation, landmark identification and safety, and you should perform those exactly as taught. This page is for the sentence that earns the extra mark and for what you do on placement: state the measurement, then state what change would have to occur before you would call it real.
Why does end range measure worse?
The study found wider angle measurement reduced device reliability across 0° to 180°. [1] Soft tissue moves, the axis of rotation migrates, and the force you apply at end range varies more than you think. It is one reason the same study recommends evaluating error separately across the range rather than quoting a single figure.
Where does this fit with outcome measures?
It is the same discipline applied to a tool you hold rather than a questionnaire you hand over. Every instrument in the outcome measures library is reported with its measurement error for exactly this reason, and several of them have a minimal important change smaller than their own detectable change — the same trap in a different form.
References
- Kiatkulanusorn S, Luangpon N, Srijunto W, et al. Analysis of the concurrent validity and reliability of five common clinical goniometric devices. Scientific Reports. 2023 Nov 27;13(1):20931. doi:10.1038/s41598-023-48344-6 PMID 38017058 Validity and reliability study, five devices
- Svensson M, Lind V, Löfgren Harringe M. Measurement of knee joint range of motion with a digital goniometer: A reliability study. Physiotherapy Research International. 2019 Apr;24(2):e1765. doi:10.1002/pri.1765 PMID 30589162 Reliability study
- van Rijn SF, Zwerus EL, Koenraadt KL, et al. The reliability and validity of goniometric elbow measurements in adults: A systematic review of the literature. Shoulder & Elbow. 2018 Oct;10(4):274–284. doi:10.1177/1758573218774326 PMID 30214494 Systematic review
- van Kooij YE, Fink A, Nijhuis-van der Sanden MW, et al. The reliability and measurement error of protractor-based goniometry of the fingers: A systematic review. Journal of Hand Therapy. 2017 Oct-Dec;30(4):457–467. doi:10.1016/j.jht.2017.02.012 PMID 28389132 Systematic review (COSMIN)
- Cools AM, De Wilde L, Van Tongel A, et al. Measuring shoulder external and internal rotation strength and range of motion: comprehensive intra-rater and inter-rater reliability study of several testing protocols. Journal of Shoulder and Elbow Surgery. 2014 Oct;23(10):1454–61. doi:10.1016/j.jse.2014.01.006 PMID 24726484 Reliability study
About this resource
- Written by
- Dr Kashina Arora (PT)BPT, MPT · Senior Physiotherapist · HCMCT Manipal Hospital, Dwarka, Delhi
- Reviewed by
- Dr Chitrakshi Sharma (PT)BPT, MPT · Head of Department · APARC Health and Motion, Janakpuri · not the author
- Chief Editor
- Dr Dharam Pandey (PT)MPT; PhD
- Evidence grade
- ModerateSee "How certain is this?"
- Last reviewed
- 16 August 2026Next review due 16 August 2028
How to use this
Written to be learned from, not memorised.
Every figure on this page is given as an error in degrees as well as a correlation, because the correlation is the number that flatters the tool and the error is the number you have to act on. Faculty may use this page in teaching with attribution. It carries its review date and its next review date, so you can see at a glance whether it is current before you put it in front of a cohort.
