Student library · Practical skills
Observational Gait Analysis
The total score behaves well — inter-rater ICC 0.91, with a proposed minimal detectable change of 6 points. The individual joints behave much worse, with per-item kappa running from 0.15 to 0.87. Which number you quote depends on what you are claiming.
In one line. Observational gait analysis is watching someone walk and scoring what you see, usually with a structured scale and increasingly from video. It is what almost every physiotherapist actually does, because instrumented three-dimensional analysis is available to almost none of them.
Done with a validated scale, from video, by a trained observer, it is defensible. Done by eye in real time and recorded as a list of joint-by-joint deviations, most of what you write down is not reproducible.
Distal is easier than proximal, consistently. Two independent studies found reliability better for the foot and knee than for the pelvis, hip and trunk. [2][3] That is not a comment on your skill; it reflects how much of a proximal segment's motion is visible from the outside and how much soft tissue moves over it. Score the ankle and knee with confidence, and treat your pelvis and trunk observations as hypotheses.
The numbers you actually need
| Question | Finding | Source and caveat |
|---|---|---|
| How reliable is the total score? | Good to excellent. Edinburgh Visual Gait Score intra-rater ICC 0.90 to 0.97 and inter-rater ICC 0.91 | Children with cerebral palsy, mostly GMFCS levels I to II with mild to moderate gait pathology; few at level III [1] |
| What change is real? | MDC90 ranged 3.6 to 6.0, and the authors propose an MDC of 6.0 for the EVGS | Same study. A change of three or four points on this scale is inside the measurement error [1] |
| And item by item? | Much weaker. Intra-observer complete agreement 65% to 98.3% and inter-observer 61.7% to 92.5%, with kappa values from 0.15 to 0.87 | Smartphone video and a motion analysis app, children with cerebral palsy. Reliability was better for distal than proximal segments [2] |
| Does training matter? | Yes. Observers trained in gait analysis produced significantly better results, and the authors conclude observers should either be used to interpreting clinical gait analysis or receive video analysis training first | Reliability also differed significantly between stance and swing phase [3] |
| Does agreement mean validity? | No. Video-based Rivermead Visual Gait Assessment showed good-to-excellent agreement between raters and between assessments (correlation 0.94 to 0.95) but correlated with the Berg Balance Scale at only r = 0.4 | Post-stroke hemiparesis. The authors describe the validity as acceptable; it is also a reminder that raters agreeing is a different question from the score meaning what you think [4] |
| Can 2D video replace 3D? | For a tightly defined measure, apparently yes. Two-dimensional measurement of peak transverse-plane upper torso rotation showed intra-rater ICC 0.989 to 0.999, inter-rater 0.990 to 0.995, MDC 0.39 to 1.4 degrees, and correlated with 3D at r of at least 0.986 | Runners, one specific variable. This does not license treating 2D video as equivalent to 3D analysis in general [5] |
| Does it track function? | Reasonably. EVGS correlated with the GMFM-66 at r = -0.69 to -0.73 | Which supports construct validity, in the population studied [1] |
Where students get this wrong
1. Quoting the total's reliability to defend an item-level claim
Inter-rater ICC for the EVGS total was 0.91. [1] Per-item kappa in another EVGS study ran as low as 0.15. [2] If you write "reduced hip extension in terminal stance", that is an item-level claim and the item-level statistic applies to it. The total is more reliable than any of its parts, which is a property of sums, not evidence that you saw the hip correctly.
2. Scoring in real time and treating it like video
Every reliability figure on this page comes from video or from structured scoring. [1][2][4] Real-time observation at walking speed gives you one pass, no replay and no slow motion. Where a decision matters, record it. A phone on a tripod is enough — one of these studies used exactly that. [2]
3. Believing you can see the pelvis
Two studies, different populations and different scales, both found proximal segments less reliable than distal. [2][3] Trunk, pelvis and hip are the segments where your confidence should be lowest and where you are most likely to be describing clothing, soft tissue or your own expectation.
4. Assuming a small change means improvement
The proposed MDC for the EVGS is 6.0 points. [1] A patient moving from 14 to 11 has not demonstrably changed. This is the same discipline as goniometry and manual muscle testing: report the change against the error, not the change alone.
5. Treating rater agreement as proof the score means something
The RVGA study is the clearest example: raters agreed at 0.94 to 0.95 with each other and the score correlated with the Berg Balance Scale at 0.4. [4] Two people can reliably measure the same thing while that thing tells you less about the patient than you assumed. Reliability and validity are separate questions and this literature reports both.
6. Generalising the 2D-versus-3D result
The near-perfect agreement between 2D and 3D applies to one variable, peak transverse-plane upper torso rotation, in runners. [5] It is a demonstration that a carefully defined 2D measure can be excellent, not evidence that phone video substitutes for a gait laboratory. Quote it for what it tested.
What the evidence supports — and what it does not
Supported
- Using a validated scale total rather than free-text observations. [1]
- Scoring from video rather than in real time. [1][2][4]
- Training before scoring. Trained observers did significantly better. [3]
- More confidence in foot and knee than in pelvis and trunk. [2][3]
- Reporting change against an MDC, proposed as 6.0 for the EVGS. [1]
Not supported
- Item-level claims backed by the total's reliability. [1][2]
- Confident proximal-segment observations. [2][3]
- Reading rater agreement as validity. r = 0.4 against Berg. [4]
- Treating 2D video as generally equivalent to 3D. [5]
- Transferring these figures beyond the populations studied — mostly children with cerebral palsy at GMFCS I to II. [1][2]
How certain is this?
Evidence grade: Low to moderate.
The direction is consistent and replicated: total scores are reliable, items are not uniformly so, distal beats proximal, and training helps. Two independent groups found the proximal-distal pattern in different populations with different scales. [2][3]
What limits the grade is population. Three of these five studies are in children with cerebral palsy, [1][2][3] and the largest EVGS reliability study explicitly notes it had few children at GMFCS level III, so its figures apply to mild to moderate gait pathology. The adult stroke [4] and running [5] studies each examine a single scale or a single variable. No study here covers the general adult musculoskeletal caseload most students will observe.
A note on the dates. Three of these papers predate 2021. They are reliability derivations for scales whose scoring has not changed, and psychometric evidence does not date the way treatment evidence does.
Common questions
Should I use a scale or just describe what I see?
Use a scale. Every reliability figure on this page comes from structured scoring, and free-text description has no established reliability at all. A scale also forces you to look at segments you would otherwise skip and gives you a total that behaves better than any single observation. [1]
Is phone video good enough?
For observational scoring, the evidence is encouraging: one of these reliability studies used smartphone video with a motion analysis application and reported intra-observer agreement up to 98.3%. [2] For precise kinematic measurement, a carefully defined 2D variable matched 3D almost perfectly in one study, [5] but that was one variable under controlled conditions. Video for scoring, yes; video as a gait laboratory, no.
How much change counts?
Six points on the EVGS, on the authors' own proposal, with MDC90 across their analyses ranging 3.6 to 6.0. [1] For other scales, look up the equivalent figure rather than assuming, and state it alongside the change you are reporting.
Why is my agreement with my supervisor poor on some joints?
Probably because those joints are hard, not because you are wrong. Per-item kappa in one EVGS study ranged from 0.15 to 0.87, [2] and reliability differed significantly between stance and swing phase in another. [3] Compare scores item by item with your supervisor and focus discussion on the segments where you disagree; that is where the learning is.
Does this apply to adults?
Partly. The strongest reliability data here are from children with cerebral palsy at GMFCS levels I to II. [1][2] The adult evidence on this page is one study of the Rivermead Visual Gait Assessment after stroke, which reported rater agreement of 0.94 to 0.95. [4] Treat the general principles as transferable and the specific thresholds as population-bound.
References
- Abe H, Koyanagi S, Kusumoto Y, et al. Intra-rater and inter-rater reliability, minimal detectable change, and construct validity of the Edinburgh Visual Gait Score in children with cerebral palsy. Gait & Posture. 2022 May;94:119–123. doi:10.1016/j.gaitpost.2022.03.004 PMID 35279565 Reliability and validity study
- Aroojis A, Sagade B, Chand S. Usability and Reliability of the Edinburgh Visual Gait Score in Children with Spastic Cerebral Palsy Using Smartphone Slow-Motion Video Technology and a Motion Analysis Application: A Pilot Study. Indian Journal of Orthopaedics. 2021 Aug;55(4):931–938. doi:10.1007/s43465-020-00332-y PMID 34194650 Reliability study
- Viehweger E, Zürcher Pfund L, Hélix M, et al. Influence of clinical and gait analysis experience on reliability of observational gait analysis (Edinburgh Gait Score Reliability). Annals of Physical and Rehabilitation Medicine. 2010 Nov;53(9):535–46. doi:10.1016/j.rehab.2010.09.002 PMID 20952267 Reliability study
- Arya KN, Pandian S, Kumar V, et al. Post-stroke Visual Gait Measure for Developing Countries: A Reliability and Validity Study. Neurology India. 2019 Jul-Aug;67(4):1033–1040. doi:10.4103/0028-3886.266273 PMID 31512628 Reliability and validity study
- Weber CF, McClinton S. VALIDITY AND RELIABILITY OF VIDEO-BASED ANALYSIS OF UPPER TRUNK ROTATION DURING RUNNING. International Journal of Sports Physical Therapy. 2020 Dec;15(6):910–919. doi:10.26603/ijspt20200910 PMID 33344007 Concurrent validity study
About this resource
- Written by
- Dr Priya Bhati (PT)BPT, MPT · Physiotherapist · ShardaCare Healthcity, Greater Noida
- Reviewed by
- Dr Ravikant Mishra (PT)BPT, MPT · Head of Department · Department of Physiotherapy & Rehabilitation Science, ShardaCare Healthcity, Greater Noida · not the author
- Chief Editor
- Dr Dharam Pandey (PT)MPT; PhD
- Evidence grade
- Low to moderateSee "How certain is this?"
- Last reviewed
- 16 August 2026Next review due 16 August 2028
How to use this
Written to be learned from, not memorised.
This page separates the reliability of a gait scale's total from the reliability of the individual observations students are actually asked to record. Faculty may use this page in teaching with attribution. It carries its review date and its next review date, so you can see at a glance whether it is current before you put it in front of a cohort.
