Skip to content

Student library · Practical skills

Manual Muscle Testing

One study of the MRC scale reports almost perfect agreement. Another, using the same scale, reports significant disagreement. Both are correct — and the difference between them is which reliability statistic the authors chose to report.

Evidence MRC inter-rater kappa 0.77 to 0.81· Same data, Fleiss' kappa 0.416· Hand-held dynamometer, two examiners: ICC 0.11 to 0.28

In one line. Manual muscle testing grades strength 0 to 5 against gravity and manual resistance. It is fast, needs no equipment, and above grade 3 the examiner's own strength becomes part of the measurement.

It remains the correct first test in most neurological and many musculoskeletal assessments. What it is not is a precise or interchangeable number, and this page is about the size of that imprecision.

The statistic changes the verdict. Given the same set of MRC gradings from 109 clinicians, the authors calculated Krippendorff's alpha at 0.788, Gwet's AC2 at 0.808, and Fleiss' kappa at 0.416. [2] Two of those read as good agreement and one as moderate at best. When you read that a test is "reliable", check which coefficient was used and on what distribution of scores — kappa is penalised when most ratings cluster in a few categories, which is exactly what MMT gradings do.

The numbers you actually need

QuestionFindingSource and caveat
Do two examiners agree on an MRC grade? Yes, on one reading. Inter-rater weighted kappa 0.77 to 0.81 and intra-rater 0.81 to 0.88 across finger extension, wrist extension and grip — described as almost perfect agreement 31 patients with radial nerve palsy. A narrow clinical population and a small sample [1]
And on another reading? No. Significant dispersion in motor grading of example vignettes, with Fleiss' kappa 0.416 (95% CI 0.413 to 0.419) 109 respondents. The authors' conclusion is that there is significant disagreement among providers applying the MRC scale [2]
Does it correlate with measured force? Moderately. Spearman correlation of 0.78 between maximal relative force and the median MRC score Same study as [1]. A correlation of 0.78 leaves substantial unexplained variance [1]
Is a hand-held dynamometer the fix? Only if the same person uses it. Traditional hand-held dynamometry had good intra-examiner reliability (ICC 0.79 to 0.93) and poor inter-examiner reliability (ICC 0.11 to 0.28) Smallest real difference was 21% to 43% for the traditional device, against 9% to 16% when the examiner's resisting force was mechanically enhanced [3]
Why does the examiner matter so much? Because above grade 3 you are measuring the patient against you. Enhancing the examiner's resisting force improved inter-examiner reliability to ICC 0.98 Same study. This is the clearest available demonstration that the limiting factor is the tester, not the patient [3]
Reliable and valid are different questions Test-retest ICC 0.96 and inter-rater ICC 0.95 to 0.97 for knee flexion and extension — but validity against a gold standard was moderate for knee flexion (ICC 0.75) and poor for knee extension (ICC 0.37, 95% CI 0 to 0.73) Healthy adults. The authors state absolute values may not be comparable to gold-standard strength assessment [4]
Can it track change over time? At group level, yes; at individual muscle level, poorly. Summed strength across four muscle groups captured change over a year (P < 0.01) and correlated with function (rho 0.68 to 0.92), but individual muscle groups showed high variability between visits The authors suggest selecting functional measures with less variability as trial endpoints rather than strength testing [5]

Where students get this wrong

1. Treating 4 and 5 as measurements

Grades 0 to 3 are anchored to observable events: no contraction, flicker, movement with gravity eliminated, movement against gravity. Grade 4 and 5 are anchored to how hard you can push. When the examiner's resisting force was mechanically standardised, inter-examiner reliability rose from ICC 0.11–0.28 to 0.98. [3] Almost all the error in the top half of the scale is the tester.

2. Comparing your grade to a colleague's

Intra-rater agreement is consistently better than inter-rater agreement, [1][3] and with a hand-held dynamometer the inter-examiner ICC fell as low as 0.11. [3] If a grade was recorded at admission by someone else, re-test it yourself before concluding the patient has changed. The same rule applies to goniometry, and for the same reason.

3. Believing a good ICC means a small error

The knee study reports test-retest ICC 0.96 alongside validity against the gold standard of 0.37 for knee extension. [4] Reliability asks whether you get the same answer twice; validity asks whether the answer is right. A test can be beautifully consistent and consistently wrong.

4. Reading "reliable" without reading the coefficient

The clearest example in this literature: Krippendorff's alpha 0.788, Gwet's AC2 0.808 and Fleiss' kappa 0.416 from the same ratings. [2] Kappa corrects for chance agreement and is punished when ratings cluster in few categories, which MMT scores do. Naming the statistic is not pedantry in your assignments; it is the difference between the two opposite conclusions on this page.

5. Tracking a single muscle over weeks

Individual muscle groups showed high variability between visits even where a summed score across four groups detected real change over a year. [5] If you want to show a patient is improving, a summed score or a functional measure will do it more honestly than "quadriceps went from 4 to 4+".

6. Assuming a small study settles it

The "almost perfect agreement" figures come from 31 patients with one specific diagnosis. [1] That is a legitimate finding about radial palsy assessment and a weak basis for a general claim about MMT. The larger survey pointing the other way had 109 raters but used vignettes rather than patients. [2] Neither is decisive, which is why this page reports both.

What the evidence supports — and what it does not

Supported

  • MMT as a fast screen for weakness, particularly across grades 0 to 3. [1]
  • The same examiner re-testing the same patient. [1][3]
  • Dynamometry where a number is needed, used by one examiner. [3][4]
  • Summed or functional scores for tracking change. [5]
  • Standardising the resisting force where equipment allows. ICC rose to 0.98. [3]

Not supported

  • Treating grades 4 and 5 as objective. [3]
  • Comparing grades between examiners. ICC as low as 0.11. [3]
  • Reading a high ICC as accuracy. Validity 0.37 for knee extension. [4]
  • Tracking a single muscle group between visits. [5]
  • Quoting "MMT is reliable" without the coefficient and population. [1][2]

How certain is this?

Evidence grade: Low to moderate.

The individual studies are competently done but small and narrow: 31 patients with radial palsy, [1] healthy adults for the knee work, [4] and vignettes rather than patients for the largest rater sample. [2] No study here examines MMT across the range of patients a physiotherapy student will actually test.

What raises confidence is that the central mechanism — that the examiner's own strength limits the measurement above grade 3 — is demonstrated directly rather than inferred, by showing that standardising the resisting force takes inter-examiner ICC from 0.11–0.28 to 0.98. [3] That is a strong result from a clean comparison.

A note on the dates. Four of these five papers predate 2021. They are the primary reliability derivations for a test whose technique has not changed, and psychometric evidence does not date the way treatment evidence does.

Common questions

Is manual muscle testing worth doing at all?

Yes. It is fast, needs nothing, and grades 0 to 3 are anchored to events you can observe rather than to your own strength. It correlated with measured force at rho 0.78 in one study. [1] The caution is about precision in the top half of the scale and about comparing across examiners, not about whether to test.

Should I use a hand-held dynamometer instead?

If you need a number and you will take both measurements yourself, yes — intra-examiner ICC was 0.79 to 0.93. [3] If someone else took the baseline, the dynamometer is not obviously better: inter-examiner ICC was 0.11 to 0.28 with smallest real difference of 21% to 43%. [3] The device does not solve the problem; standardising the resisting force does.

Why do the two reliability studies disagree?

Different populations, different tasks, and above all different statistics. One used weighted kappa on real patients with one diagnosis and found 0.77 to 0.81. [1] The other used vignettes across 109 clinicians and reported three coefficients on the same data: 0.788, 0.808 and 0.416. [2] The last is Fleiss' kappa, which is penalised when ratings cluster in a few categories. Report the coefficient, not just the adjective.

My patient improved from 4 to 4+. Is that real?

Treat it as unproven. Individual muscle groups showed high variability between visits even in a study where a summed score detected genuine change over a year. [5] If it matters clinically, corroborate it with something less examiner-dependent — a functional task, a timed test, or a dynamometer reading you took yourself both times.

How does this connect to outcome measures?

Directly. The instruments in the outcome measures library are reported with their measurement error for the same reason this page reports smallest real difference: a change smaller than the error is not a finding. MMT is the version of that problem you hold in your hands.

References

  1. Paternostro-Sluga T, Grim-Stieger M, Posch M, et al. Reliability and validity of the Medical Research Council (MRC) scale and a modified scale for testing muscle strength in patients with radial palsy. Journal of Rehabilitation Medicine. 2008 Aug;40(8):665–71. doi:10.2340/16501977-0235 PMID 19020701 Reliability and validity study
  2. Smith BW, Sakamuri S, Flavin KE, et al. Assessment of variability in motor grading and patient-reported outcome reporting: a multi-specialty, multi-national survey. Acta Neurochirurgica. 2021 Jul;163(7):2077–2087. doi:10.1007/s00701-021-04861-9 PMID 33990886 Survey of rater agreement
  3. Lu TW, Hsu HC, Chang LY, et al. Enhancing the examiner's resisting force improves the reliability of manual muscle strength measurements: comparison of a new device with hand-held dynamometry. Journal of Rehabilitation Medicine. 2007 Nov;39(9):679–84. doi:10.2340/16501977-0107 PMID 17999004 Reliability study
  4. Neil SE, Myring A, Peeters MJ, et al. Reliability and validity of the Performance Recorder 1 for measuring isometric knee flexor and extensor strength. Physiotherapy Theory and Practice. 2013 Nov;29(8):639–47. doi:10.3109/09593985.2013.779337 PMID 23724831 Reliability and validity study
  5. Reash NF, James MK, Alfano LN, et al. Comparison of strength testing modalities in dysferlinopathy. Muscle & Nerve. 2022 Aug;66(2):159–166. doi:10.1002/mus.27570 PMID 35506767 Longitudinal responsiveness study

About this resource

How to use this

Written to be learned from, not memorised.

This page reports two studies that reach opposite conclusions about the same scale, and explains why, because the explanation is the thing worth learning. Faculty may use this page in teaching with attribution. It carries its review date and its next review date, so you can see at a glance whether it is current before you put it in front of a cohort.