Clinical library · Clinical reasoning
Yellow Flags and the STarT Back Tool
Psychosocial obstacles to recovery are real and worth identifying. Whether a nine-item questionnaire can sort patients into risk groups accurately enough to drive treatment is a separate question, and the answer depends heavily on who you are asking it about.
In one line. "Yellow flags" are psychological factors that predict persistent disability. The STarT Back Tool is a nine-item questionnaire that sorts patients with low back pain into low, medium and high risk groups so that treatment intensity can be matched to prognosis.
The colour system is wider than most clinicians use it. A systematic review of 15 guidelines describes yellow flags (psychological factors), blue flags (perceptions about the relationship between work and health), black flags (system or contextual obstacles) and orange flags (psychiatric symptoms). Guideline recommendations almost exclusively concerned yellow or black flags. [1]
Two claims, not one. "Psychosocial factors matter in low back pain" and "the STarT Back Tool accurately predicts who will do badly" are different statements with different evidence. The first is widely accepted. The second is contested, and the trial results across four countries have been contrasting. [3] This page keeps them apart.
What the evidence shows
| Question | Finding | Source and quality |
|---|---|---|
| Does stratified care work? | Yes in the original trial. Adjusted mean change in Roland-Morris score was higher with stratified care at 4 months (4.7 vs 3.0, difference 1.81, 95% CI 1.06 to 2.57) and at 12 months (4.3 vs 3.3, difference 1.06, 95% CI 0.25 to 1.86) | 851 patients randomised, 568 intervention and 283 control. Effect sizes were 0.32 (0.19 to 0.45) at 4 months and 0.19 (0.04 to 0.33) at 12 months — small [2] |
| Was it cost-effective? | At 12 months, 0.039 additional quality-adjusted life years and lower cost (£240.01 vs £274.40) | Same trial; UK primary care, 2011 [2] |
| Has it replicated? | Not consistently. At least 7 randomised trials published between 2011 and 2023 across 4 countries, with contrasting results — some showing effectiveness or efficiency, others finding no benefit | Review by the tool's own developers. Concludes the tool's predictive performance cannot alone explain the variation; implementation challenges largely do [3] |
| Does it predict outcome in older adults? | Poorly. AUC 0.65, 0.67 and 0.65 at 3, 6 and 12 months. Accuracy measures were poor at all time points, with particularly poor sensitivity and negative likelihood ratios | 452 primary care patients aged 55 or over. The authors' conclusion is that predictive validity in this group was poor and the tool may need recalibration before widespread use in older adults [4] |
| Do medium and high risk groups actually differ? | Not significantly, in older adults. No statistically significant differences in odds between medium and high risk at any time point | Medium risk n = 118, high risk n = 27. The high-risk confidence intervals are very wide (8.90, 95% CI 1.83 to 43.24 at 3 months) [4] |
| When should it be administered? | Later may beat baseline. In an emergency department cohort, the tool at 6 weeks predicted 3-month disability more accurately than at baseline | 118 enrolled, 67 (56.7%) completed follow-up. 54 of those (80.6%) reported more than 30% improvement at 3 months, so most recovered regardless [5] |
| Do guidelines tell you what to do about yellow flags? | Not clearly. Psychosocial factors were addressed in assessment by 13 of 15 guidelines and in management by 14 of 15, but recommendations varied widely and lacked detail | The supporting evidence was generally of very low quality; the review concludes guidelines do not give clinicians clear instructions [1] |
Where it misleads
1. The original effect was real but small
The between-group difference in Roland-Morris score was 1.81 points at 4 months and 1.06 at 12 months, with effect sizes of 0.32 and 0.19. [2] That is a small effect that faded over the year. It was achieved at lower cost with a small QALY gain, which is what made it interesting at a health-system level — but it is not the transformative individual-patient result the tool is sometimes presented as delivering.
2. Replication has been inconsistent, and the developers say so
Seven or more randomised trials across four countries, with some finding effectiveness and others finding no benefit over comparison interventions. [3] The review reaching that conclusion is authored by the group that developed the tool, which makes the negative finding harder to dismiss rather than easier. Their reading is that implementation, not prediction, explains most of the variation — but that is itself a caution: a tool that only works when the whole system around it changes is not a questionnaire, it is a service redesign.
3. It was not built for, and does not perform well in, older adults
In 452 primary care patients aged 55 or over, area under the curve ran 0.65 to 0.67 and accuracy measures were poor at every time point. [4] Sensitivity and the negative likelihood ratio were singled out as particularly poor, which means a low-risk result in an older adult does relatively little to rule out a poor outcome. Medium and high risk groups did not differ significantly from each other at any time point.
4. Most people get better anyway, which flatters any prognostic tool
In the emergency department cohort, 80.6% of those followed up reported more than 30% improvement at 3 months. [5] When the base rate of recovery is that high, a tool can look accurate simply by predicting recovery. This is why sensitivity and negative likelihood ratios matter more than headline accuracy, and it is exactly where the older-adult data were weakest. [4]
5. "Screen for yellow flags" is not an intervention
Guidelines address psychosocial factors almost universally in principle — 13 of 15 for assessment, 14 of 15 for management — but the recommendations varied widely, lacked detail, and rested on evidence of generally very low quality. [1] Identifying a high-risk patient is only useful if there is a different and better thing to do for them. The developers' own recommendation for future research includes developing better treatments for high-risk patients, which concedes the point. [3]
6. Timing changes the answer
Administering the tool at 6 weeks predicted 3-month disability better than administering it at baseline in the emergency department cohort. [5] Early scores partly capture acute distress that resolves on its own. A high score in week one and a high score in week six are not the same signal.
What the evidence supports — and what it does not
Supported
- Assessing psychosocial obstacles to recovery as part of routine low back pain care. [1]
- Stratified care producing a small benefit at lower cost in UK primary care. [2]
- Using the tool to structure a conversation about fear, distress and expectations, rather than as a verdict.
- Re-administering it after a few weeks rather than relying on the first-visit score. [5]
Not supported
- Treating the risk group as a reliable individual prognosis, particularly in older adults. [4]
- Assuming the original trial result generalises. Seven trials, four countries, contrasting results. [3]
- Distinguishing medium from high risk in adults aged 55 and over. [4]
- Screening alone changing outcomes without a matched pathway behind it. [3]
- Any single guideline's yellow-flag protocol as the standard. Recommendations vary widely on very low quality evidence. [1]
How certain is this?
Evidence grade: Moderate.
The underlying trial is large, randomised, and well reported, with confidence intervals that exclude no effect at both time points. [2] What lowers the grade is everything after it: replication across four countries has been inconsistent, [3] and predictive validity in a population the tool was not developed for is poor. [4]
The guideline review is systematic and preregistered, and its central finding — that recommendations vary widely and rest on very low quality evidence — is a statement about the field, not about one study. [1]
The emergency department study is small (118 enrolled, 67 completing) and single-setting, so its timing finding is suggestive rather than established. [5]
What would change the grade: predictive validity studies in the populations where the tool is actually being used outside UK primary care, and trials of treatments specifically designed for the high-risk group rather than more of the same at higher intensity.
Common questions
Should I still use the STarT Back Tool?
In adult primary care low back pain, with a genuine matched pathway behind it, yes — that is the setting where it produced a small benefit at lower cost. [2] Outside that setting be cautious. Its predictive validity in adults aged 55 and over was poor, [4] and across seven trials in four countries results have been contrasting. [3] Use it to structure assessment, not to issue a prognosis.
What is the difference between yellow, blue, black and orange flags?
As described in the guideline review: yellow flags are psychological factors, blue flags are perceptions about the relationship between work and health, black flags are system or contextual obstacles, and orange flags are psychiatric symptoms. [1] In practice, guideline recommendations almost exclusively concerned yellow or black flags, so the other two categories are much less developed.
Why did the tool work in one trial and not in others?
The developers addressed this directly. Their conclusion is that although there is room for improving the tool's predictive value, its performance in allocating people to risk categories cannot alone explain the variation — the differences are largely explained by how successfully the whole stratified-care approach was implemented and by the difficulty of changing professional practice. [3]
Does it work in older patients?
Not well. In 452 primary care patients aged 55 or over, AUC was 0.65 to 0.67 across 3, 6 and 12 months and accuracy measures were poor throughout, with sensitivity and negative likelihood ratios singled out. Medium and high risk groups did not differ significantly at any time point. The authors suggest recalibration or extension before widespread use in this group. [4]
When should I administer it?
Not necessarily at the first visit. In an emergency department cohort, the score at 6 weeks predicted 3-month disability more accurately than the baseline score. [5] Early distress often settles on its own — 80.6% of those followed up improved by more than 30% at 3 months regardless.
What do I do with a high-risk result?
Honestly, the evidence is thinner here than for identifying the patient. Guidelines address management of psychosocial factors almost universally but with recommendations that vary widely and rest on very low quality evidence, [1] and the tool's own developers list "developing better treatments for patients at high risk of poor outcomes" as an open research priority. [3] Address what the score surfaced — fear, distress, work obstacles — and see low back pain.
References
- Knoop J, Rutten G, Lever C, et al. Lack of Consensus Across Clinical Guidelines Regarding the Role of Psychosocial Factors Within Low Back Pain Care: A Systematic Review. The Journal of Pain. 2021 Dec;22(12):1545–1559. doi:10.1016/j.jpain.2021.04.013 PMID 34033963 Systematic review of clinical practice guidelines
- Hill JC, Whitehurst DG, Lewis M, et al. Comparison of stratified primary care management for low back pain with current best practice (STarT Back): a randomised controlled trial. The Lancet. 2011 Oct 29;378(9802):1560–71. doi:10.1016/S0140-6736(11)60937-9 PMID 21963002 Randomised controlled trial
- Croft P, Hill JC, Foster NE, et al. Stratified health care for low back pain using the STarT Back approach: holy grail or doomed to fail?. Pain. 2024 Dec 1;165(12):2679–2692. doi:10.1097/j.pain.0000000000003319 PMID 39037849 Review of seven randomised trials
- Vigdal ØN, Flugstad S, Storheim K, et al. Predictive validity of the STarT Back screening tool among older adults with back pain. European Journal of Pain. 2024 Oct;28(9):1559–1570. doi:10.1002/ejp.2281 PMID 38752601 Prospective cohort study
- Treanor C, Brogan S, Burke Y, et al. Prospective observational study investigating the predictive validity of the STarT Back tool and the clinical effectiveness of stratified care in an emergency department setting. European Spine Journal. 2022 Nov;31(11):2866–2874. doi:10.1007/s00586-022-07264-1 PMID 35786771 Prospective observational study
About this resource
- Written by
- Dr Arjit Vashishtha (PT)BPT, MPT · Head of Department · Manipal Hospital, Ghaziabad, Uttar Pradesh
- Reviewed by
- Dr Anuj Mishra (PT)BPT, MPT · Head of Department · Department of Physiotherapy & Rehabilitation Science, Shanti Mukand Hospital, Karkardooma, Delhi · not the author
- Chief Editor
- Dr Dharam Pandey (PT)MPT; PhD
- Evidence grade
- ModerateSee "How certain is this?"
- Last reviewed
- 16 August 2026Next review due 16 August 2028
Using this in clinic
Every figure here is traceable to its source.
This page separates the case for assessing psychosocial obstacles, which is strong, from the predictive accuracy of the questionnaire, which depends on the population. Where a value could not be verified against the paper it came from, it is not on this page, and the omission is stated rather than filled with a number from a secondary source.
