Skip to content

Clinical library · Outcome measures

Neck Disability Index

The standard self-report measure for neck pain, with good reliability and a minimal clinically important difference that has been reported anywhere from 5 to 33 percent depending on the population and the method.

Evidence 10 sections · 0–50 (0–100%)· Meta-analysis of 79 studies· Pooled test-retest ICC 0.91

In one line. Ten sections modelled on the Oswestry — pain intensity, personal care, lifting, reading, headaches, concentration, work, driving, sleeping and recreation — each scored 0 to 5, giving 0 to 50, often reported as a percentage.

It is the most used neck-specific measure and its reliability is well established. The interpretive problem is the same one that affects most patient-reported measures, and here it is unusually wide: the published important-change threshold spans a sixfold range.

Which threshold? The 2024 meta-analysis of 79 studies found minimal detectable change reported from 3% to 27% and minimal clinically important difference from 5% to 33%, and concludes that both were around 15% — that is, 7.5 of 50 points. [1] An earlier review put the minimal detectable change at around 5/50 for uncomplicated neck pain and up to 10/50 for cervical radiculopathy, with clinically important difference reported inconsistently from 5/50 to 19/50. [2] The population matters more than the average.

The numbers you actually need

PropertyValueSource and caveat
Test-retest reliability Pooled ICC 0.91 (95% CI 0.90 to 0.93) 79 studies, 70 at low risk of bias [1]. An earlier review reported ICCs ranging 0.50 to 0.98, influenced by test interval and definition of "stable" [2]
Internal consistency Alpha above 0.81 Same meta-analysis [1]
Minimal detectable change Reported 3% to 27%; about 15% overall Around 5/50 for uncomplicated neck pain, up to 10/50 for cervical radiculopathy [1][2]
MCID Reported 5% to 33%; about 15% (7.5/50) overall [1]. A separate cohort gave an MDC of 10.5/50 with an ROC-optimal cut-off of 3.5 [3]
Minimal important change, registry 17 points at 6 months; 9 at 12 months 551 patients with neck pain. Anchor-based calculation was more accurate than distribution-based, which fell below measurement error [4]
Patient-determined important change 3.5 points, against statistical error of 2.16 Only 42 participants; 4 of 5 credibility criteria met. Lower than researcher-defined thresholds [6]
Discrimination Pooled area under the curve 0.74 (95% CI 0.68 to 0.80) Most studies detected no floor or ceiling effect [1]

What it measures

Self-reported activity limitation attributed to neck pain, including two items — headaches and concentration — that are not present in its lumbar ancestor and that matter clinically in cervicogenic headache and whiplash-associated disorder.

Dimensionality is contested in a way worth knowing: of the studies examined, 13 found the index unidimensional and 15 found it two- or three-dimensional, and the reviewers conclude it can be considered unidimensional in most situations. [1] An earlier review reported it as one-dimensional and interpretable as an interval scale, while noting some studies question those assumptions. [2]

Where it misleads

1. The important-change threshold depends on who defined "important"

Three approaches give three answers in the same instrument. Researcher-defined anchor methods in a registry of 551 patients gave a minimal important change of 17 points at six months and 9 at twelve. [4] The pooled literature centres on about 15%, or 7.5/50. [1] Asking patients directly which changes mattered to them gave 3.5 points — barely above the statistical error of 2.16 — and the authors note that patient-determined importance may yield a lower threshold than researcher-defined methods. [6] None is wrong; they answer different questions.

2. Cervical radiculopathy needs a higher bar than mechanical neck pain

Minimal detectable change is around 5/50 in uncomplicated neck pain but up to 10/50 in cervical radiculopathy. [2] Using the general figure in a radicular population will call noise a response. See cervical radiculopathy.

3. Distribution-based thresholds can fall below the measurement error

In the registry study, minimal important change calculated by anchor-based methods was more accurate, while distribution-based calculations fell below the measurement error for the instrument. [4] A threshold smaller than your own measurement error is not a threshold.

4. Rasch-corrected versions are not interchangeable with the original

Two Rasch-approved alternatives exist, of 8 and 5 items, which provide interval-level scoring that the ordinal 10-item version does not. In 201 patients, the mean difference between the 10-item and 5-item versions was about 10% of the total score (-4.6 points), with limits of agreement from -14.9 to 5.8. The authors state plainly that because of the size and unpredictable nature of the bias, the versions should not be used interchangeably. [5]

What the evidence supports — and what it does not

Supported

  • Good internal consistency and test-retest reliability, pooled ICC 0.91. [1]
  • Absence of floor and ceiling effects in most studies. [1]
  • Treating it as unidimensional in most situations. [1][2]
  • Anchor-based thresholds in preference to distribution-based ones. [4]
  • Population-specific measurement error — higher in cervical radiculopathy. [2]

Not supported

  • A single MCID. Reported from 5% to 33%. [1]
  • Distribution-based thresholds that fall below measurement error. [4]
  • Swapping between the 10-, 8- and 5-item versions. Bias is large and unpredictable. [5]
  • Applying uncomplicated-neck-pain thresholds to radiculopathy. [2]
  • Reading the 3.5-point patient-determined value as established. 42 participants, exploratory. [6]

How certain is this?

Evidence grade: Moderate.

The reliability evidence is strong: a meta-analysis of 79 studies, 70 of them judged at low risk of systematic bias, with a tight pooled confidence interval. [1] That is a better evidence base than most instruments on this site enjoy.

The interpretability evidence is much weaker, and its weakness is systematic rather than random — the threshold genuinely depends on population, method and whose judgement of importance is used. The 2009 review made the same observation about inconsistent clinically important differences and called for more studies in different clinical populations; [2] fifteen years later the newer meta-analysis still reports a range from 5% to 33%. [1]

The patient-determined change study is explicitly exploratory with 42 participants, and the authors applied a formal credibility instrument, meeting 4 of 5 criteria. [6] The version-agreement study is a single sample of 201 patients. [5]

What would change the grade: MCID derivation stratified by diagnosis using a pre-specified anchor, and clarity on which version a service is using.

Common questions

What change should I call meaningful?

About 15%, or 7.5 of 50 points, is the best single summary from the largest meta-analysis — while noting reported values run from 5% to 33%. [1] In uncomplicated neck pain the measurement error is around 5/50; in cervical radiculopathy up to 10/50, so the bar should be higher there. [2] Prefer anchor-based thresholds. [4]

Why is my patient's "important" improvement smaller than the published MCID?

Because published thresholds are mostly researcher-defined. When patients themselves were asked to classify their change as important or not, the patient-determined value was 3.5 points — above the statistical error of 2.16, but well below typical researcher-derived figures. The authors suggest this may be more responsive to patient-centric change. [6] It is exploratory, with 42 participants.

Is the short version equivalent?

No. Two Rasch-approved versions exist (8 and 5 items) that provide interval-level scoring the ordinal original does not, but the mean difference between the 10-item and 5-item versions was around 10% of the total score, with limits of agreement from -14.9 to 5.8. The authors state they should not be used interchangeably. [5]

Is it reliable enough for individual decisions?

Reliability is good — pooled test-retest ICC 0.91 across 79 studies, alpha above 0.81, no floor or ceiling effect in most studies. [1] Discrimination is more modest, with a pooled area under the curve of 0.74. Note the older review found ICCs ranging from 0.50 to 0.98 depending on test interval and how "stable" was defined, [2] so the pooled figure describes the literature, not your clinic.

Does it work alongside a pain score?

Yes, and the two are usually reported together. In a cohort study, the minimal detectable change was 10.5 points for the index and 4.3 for the numerical rating scale, with ROC-optimal cut-offs of 3.5 and 2.5 respectively. [3] In a registry of 551 patients both were responsive at six and twelve months. [4] See neck pain for the treatment evidence.

References

  1. Saltychev M, Pylkäs K, Karklins A, et al. Psychometric properties of neck disability index - a systematic review and meta-analysis. Disability and Rehabilitation. 2024 Nov;46(23):5415–5431. doi:10.1080/09638288.2024.2304644 PMID 38240027 Systematic review and meta-analysis
  2. MacDermid JC, Walton DM, Avery S, et al. Measurement properties of the neck disability index: a systematic review. Journal of Orthopaedic & Sports Physical Therapy. 2009 May;39(5):400–17. doi:10.2519/jospt.2009.2930 PMID 19521015 Systematic review
  3. Pool JJ, Ostelo RW, Hoving JL, et al. Minimal clinically important change of the Neck Disability Index and the Numerical Rating Scale for patients with neck pain. Spine. 2007 Dec 15;32(26):3047–51. doi:10.1097/BRS.0b013e31815cf75b PMID 18091500 Prospective cohort study
  4. John B, Røe C, Brox JI, et al. Responsiveness and minimal important change of neck disability index and numeric pain rating scale for neck patients in the Norwegian neck and back register. European Spine Journal. 2025 Jun;34(6):2219–2226. doi:10.1007/s00586-025-08836-7 PMID 40272496 Registry-based responsiveness study
  5. Lu Z, MacDermid JC, Nazari G. Agreement between original and Rasch-approved neck disability index. BMC Medical Research Methodology. 2020 Jul 3;20(1):180. doi:10.1186/s12874-020-01069-w PMID 32620096 Agreement study
  6. Young BA, Boland DM, Koppenhaver SL, et al. Patient-Determined Important Change for the Neck Disability Index With Application of Credibility Analysis: An Exploratory Study. Journal of Manipulative and Physiological Therapeutics. 2024 Jan-Jun;47(1-4):77–84. doi:10.1016/j.jmpt.2024.08.016 PMID 39412451 Exploratory study

About this resource

Using this in clinic

Every figure here is traceable to its source.

Every threshold on this page names the population it applies to and whose definition of importance produced it, because those two things differ by more than the measurement error does. Where a value could not be verified against the paper it came from, it is not on this page, and the omission is stated rather than filled with a number from a secondary source.