Clinical library · Outcome measures
Timed Up and Go
One of the most used tests in physiotherapy, and one of the most misused. It is a sound measure of functional mobility. As a falls screening tool — which is how most clinics use it — two meta-analyses say it does not work.
In one line. Time, in seconds, to stand from a chair, walk three metres, turn, walk back and sit down. One piece of equipment, one number, under a minute.
Its simplicity is why it spread, and why it is over-interpreted. The test measures what it measures well. The 13.5-second falls cut-off that appears on assessment forms across the world does not survive contact with the meta-analytic evidence.
The headline finding. Pooling 10 studies at the widely used threshold of 13.5 seconds, the Timed Up and Go had pooled specificity of 0.74 (95% CI 0.52 to 0.88) but pooled sensitivity of only 0.31 (95% CI 0.13 to 0.57), and logistic regression indicated the score is not a significant predictor of falls (OR 1.01, 95% CI 1.00 to 1.02, p = 0.05). The authors conclude it should not be used in isolation to identify high falls risk in community-dwelling older adults. [5]
The numbers you actually need
| Property | Value | Source and caveat |
|---|---|---|
| Reference value, 60 years and over | 9.4 seconds (95% CI 8.9 to 9.9) | Meta-analysis of 21 studies of apparently healthy elders [1] |
| By decade | 60–69: 8.1 s (7.1 to 9.0) 70–79: 9.2 s (8.2 to 10.2) 80–99: 11.3 s (10.0 to 12.7) |
Same meta-analysis; data were more homogeneous once split by age. Performance beyond the upper confidence limit is worse than average [1] |
| Adults aged 20 to 59 | Times differed significantly across decades | 200 participants, 50 per decade; the 50s were slower than the 20s, 30s and 40s. Slower times went with lower socioeconomic status, higher BMI, more comorbidities and worse perceived health [2] |
| Reliability, stroke | Excellent intra- and inter-rater | Three studies reporting ICC above 0.95, within a review of 13 studies [4] |
| Reliability, other populations | Excellent | 77 articles: excellent in typical adults, cerebral palsy, multiple sclerosis, Huntington's disease, stroke and spinal cord injury; predictive validity data limited [3] |
| Falls discrimination | Mean difference 0.63 s (high-functioning) to 3.59 s (institutional) | 53 studies, 12,832 participants. Diagnostic accuracy poor to moderate, and no cut-point can be recommended [6] |
| The 13.5-second threshold | Sensitivity 0.31, specificity 0.74 | 10 studies pooled; better at ruling in than ruling out, and not a significant predictor overall [5] |
What it measures
A sequence rather than a single ability: rising from a chair, gait initiation, steady-state walking, a turn, deceleration, and controlled sitting. That is why it correlates with so many things and why a single number cannot tell you which component failed. A patient can be slow because of quadriceps weakness at the sit-to-stand, because of a cautious turn, or because of gait speed, and the test does not distinguish them.
The equipment is a standard chair with arms, a three-metre walkway and a stopwatch. The detail that most affects comparability is the instruction: at comfortable pace or as fast as safely possible. The stroke review found wide variations in procedures and instructions across studies, and recommends they be described more clearly. [4] If you do not record which you used, your own repeat measure is not comparable.
Where it misleads
1. The falls cut-off is the biggest problem in the literature
Two independent meta-analyses reach the same conclusion by different routes. One pooled 10 studies at 13.5 seconds and found sensitivity of 0.31 — the test misses roughly seven in ten future fallers — concluding it should not be used in isolation. [5] The other, covering 53 studies and 12,832 participants, found that the majority of studies did not retain the test in multivariate analysis, that derived cut-points varied greatly, and that no cut-point can be recommended. [6]
A screening test with 31% sensitivity does not screen. If falls prevention is the goal, the evidence for the intervention is far stronger than the evidence for this particular gate to it — see falls prevention in older adults.
2. It works better in the people who need it least clearly
The pooled difference between fallers and non-fallers depended on the group: 0.63 seconds in high-functioning older people against 3.59 seconds in institutional settings. [6] In other words the test discriminates where the clinical question is often already answered by observation, and fails where a screening tool would actually add something.
3. "Excellent reliability" is not "useful prediction"
Reliability across populations is genuinely excellent [3][4] and this is routinely offered as though it settled the matter. It does not. A test can be perfectly reproducible and still predict nothing; that is precisely the pattern here, and the review of 77 articles says as much when it notes that predictive validity data were limited. [3]
4. Protocol drift makes serial measurement unreliable
Chair height, use of armrests, footwear, walking aid, and the instruction given all change the number. The stroke review's central methodological criticism is the variation in procedures and instructions across studies. [4] The same variation inside one clinic between one visit and the next produces change scores that are artefacts.
What the evidence supports — and what it does not
Supported
- Measuring basic functional mobility, and tracking it over time with a fixed protocol. [3][4]
- Comparison against age-banded reference values. [1][2]
- Use after stroke in patients able to walk, where the review recommends it for basic mobility skills. [4]
- Excellent reliability across a wide range of populations. [3][4]
Not supported
- Falls screening in isolation. Sensitivity 0.31 at the common threshold. [5]
- The 13.5-second cut-off, or any cut-off. No cut-point can be recommended. [6]
- Predicting falls in high-functioning older people. [6]
- Comparing scores across settings with different protocols. [4]
- Identifying why a patient is slow. The test does not separate its own components.
How certain is this?
Evidence grade: Moderate to high, for the negative finding.
It is unusual for the most confident statement on a page to be about what a test cannot do, but that is the position here. Two meta-analyses, one of 53 studies and 12,832 participants [6] and one pooling 10 diagnostic accuracy studies assessed with QUADAS-2, [5] agree that the test does not predict falls well enough to be used alone. Consistency across independent reviews with different methods is the strongest signal available short of a trial.
The reference values are older but remain the standard consolidation, drawn from 21 studies, and the authors themselves note clear differences between the contributing studies. [1] The normative data for adults under 60 come from a single study of 200 people at one institution and should be treated as indicative. [2]
The reliability evidence is strong but narrow in what it establishes: reproducibility, not usefulness. [3][4]
What would change the grade: a large prospective cohort deriving and validating a population-specific cut-point with adequate sensitivity, which does not currently exist.
Common questions
Should I stop using 13.5 seconds as a falls cut-off?
As a sole criterion, yes. At that threshold pooled sensitivity was 0.31 and pooled specificity 0.74, and the score was not a significant predictor of falls (OR 1.01, 95% CI 1.00 to 1.02). [5] A second meta-analysis of 53 studies concluded that no cut-point can be recommended. [6] Keep the test for mobility; use a multifactorial assessment for falls risk.
What is a normal time?
For people aged 60 and over, 9.4 seconds (95% CI 8.9 to 9.9). By decade: 8.1 seconds for 60 to 69, 9.2 for 70 to 79, and 11.3 for 80 to 99. A patient exceeding the upper confidence limit for their band can be considered to have worse than average performance. [1] For adults aged 20 to 59, normative values exist but come from a single 200-person study. [2]
Comfortable pace or fast as possible?
Either, provided you record which and never change it between visits. Variation in procedures and instructions is the main methodological criticism levelled at the test in the stroke literature. [4] The comparison that matters clinically is the patient against themselves.
Can I use it after stroke?
Yes, in patients who can walk. A systematic review of 13 studies found good convergent validity, excellent intra- and inter-rater reliability with ICCs above 0.95, and sensitivity to change, and recommends it for measuring basic mobility skills after stroke. Note that the three studies asking whether it predicts falls after stroke showed inconclusive results. [4]
Is it valid in neurological populations generally?
Reliability is excellent across typical adults and in cerebral palsy, multiple sclerosis, Huntington's disease, stroke and spinal cord injury, with strong concurrent validity in stroke and spinal cord injury. The same review notes that predictive validity data were limited and that more research is needed for young to middle-aged adults. [3]
References
- Bohannon RW. Reference values for the timed up and go test: a descriptive meta-analysis. Journal of Geriatric Physical Therapy. 2006;29(2):64–8. doi:10.1519/00139143-200608000-00004 PMID 16914068 Descriptive meta-analysis
- Kear BM, Guck TP, McGaha AL. Timed Up and Go (TUG) Test: Normative Reference Values for Ages 20 to 59 Years and Relationships With Physical and Mental Health Risk Factors. Journal of Primary Care & Community Health. 2017 Jan;8(1):9–13. doi:10.1177/2150131916659282 PMID 27450179 Cross-sectional normative study
- Christopher A, Kraft E, Olenick H, et al. The reliability and validity of the Timed Up and Go as a clinical tool in individuals with and without disabilities across a lifespan: a systematic review. Disability and Rehabilitation. 2021 Jun;43(13):1799–1813. doi:10.1080/09638288.2019.1682066 PMID 31656104 Systematic review
- Hafsteinsdóttir TB, Rensink M, Schuurmans M. Clinimetric properties of the Timed Up and Go Test for patients with stroke: a systematic review. Topics in Stroke Rehabilitation. 2014 May-Jun;21(3):197–210. doi:10.1310/tsr2103-197 PMID 24985387 Systematic review
- Barry E, Galvin R, Keogh C, et al. Is the Timed Up and Go test a useful predictor of risk of falls in community dwelling older adults: a systematic review and meta-analysis. BMC Geriatrics. 2014 Feb 1;14:14. doi:10.1186/1471-2318-14-14 PMID 24484314 Systematic review and meta-analysis
- Schoene D, Wu SM, Mikolaizak AS, et al. Discriminative ability and predictive validity of the timed up and go test in identifying older people who fall: systematic review and meta-analysis. Journal of the American Geriatrics Society. 2013 Feb;61(2):202–8. doi:10.1111/jgs.12106 PMID 23350947 Systematic review and meta-analysis
About this resource
- Written by
- Dr Dharam Pandey (PT)MPT; PhD · Chief Editor · Director & Head of Department · Department of Physiotherapy & Rehabilitation Science
- Reviewed by
- Independent external peer reviewerAnonymous third-party review · not the author
- Evidence grade
- ModerateSee "How certain is this?"
- Last reviewed
- 16 August 2026Next review due 16 August 2028
Using this in clinic
Every figure here is traceable to its source.
Every reference value, sensitivity and specificity on this page was read off the paper that reported it, with the population and the threshold stated alongside it. Where a value could not be verified against the paper it came from, it is not on this page, and the omission is stated rather than filled with a number from a secondary source.
