Skip to content

Student library · OSCE and viva

OSCE Preparation

In one analysis of communication stations, variance attributable to the examiner was more than four times the variance attributable to the student. Knowing that does not lower the bar. It tells you exactly where to spend your preparation.

Evidence Examiner variance 0.86 vs student 0.20· Reliability across stations: alpha 0.66· 17 stations needed for a sound pass or fail

In one line. The OSCE samples your performance across several short stations with different examiners. Its reliability comes almost entirely from that sampling — from how many stations there are, not from how carefully any one is marked.

The practical consequence for you: a single bad station matters less than students fear, and consistency across all of them matters more than brilliance in one.

Read this as strategy, not as an excuse. The evidence below shows that a large share of your score reflects which examiner you drew and which station you were given. None of that is under your control, so stop trying to control it. What is under your control is being reliably competent at every station rather than outstanding at one, because the exam is a sampling instrument and it rewards consistency.

What the evidence shows

QuestionFindingSource and quality
How reliable is an OSCE overall? Modest. Summary alpha across stations was 0.66 (95% CI 0.62 to 0.70), and within stations across items 0.78 (0.73 to 0.82) Meta-analysis of 188 alpha values from 39 studies. The authors' conclusion is that overall OSCE scores "are often not very reliable" [1]
What actually improves it? More stations and more examiners per station Same meta-analysis. Some OSCEs were more reliable than others for reasons not yet fully understood [1]
How much of the score is the student? Less than you would hope. Variance owing to students was 6.3%; variance owing to student-station interaction plus residual error was 66.3%; station difficulty accounted for 27.4% 463 students, generalisability analysis of a dental OSCE [3]
How much is the examiner? Examiner stringency variance was 0.86 against examinee variance of 0.20 — more than four times larger. About 11% of students' pass or fail status changed when scores were adjusted for examiner effects Communication stations rated with a 28-item guide, analysed with a multifaceted Rasch model. All facet scores were themselves reliable (0.87 to 0.99) [2]
How many stations does a sound decision need? A minimum of 12 for reliable relative decisions and 17 for reliable pass-fail decisions in that examination Same generalisability study. The authors conclude that wide sampling of stations is "at the heart" of obtaining reliable OSCE scores [3]
Are some skills harder to assess than others? Yes. Interpersonal skills were evaluated less reliably across stations, and more reliably within a station, than clinical skills Communication does not generalise across situations the way a discrete technique does [1]
Can a well-designed OSCE be reliable? Yes. G-coefficients of 0.624 to 0.823 were achieved under a standard six-station, two-rater, six-item setting, and a validated station tool reached internal and inter-observer reliability above 0.80 The first study still concludes further improvement is needed for a very high-stakes examination [4]; the second found a three-factor structure accounting for 37.1% of variance — process, communication, safety [5]

Where students get this wrong

1. Catastrophising one bad station

Station difficulty accounted for 27.4% of variance and student-station interaction plus error for 66.3%, against 6.3% for students themselves. [3] A single station going badly is substantially a property of that station and that pairing, not a verdict on you. The exam is designed on the assumption that you will have a bad station; that is why there are several.

2. Trying to work out what the examiner wants

Examiner stringency varied more than four times as much as student ability, and adjusting for it flipped the outcome for about 11% of candidates. [2] You cannot read your way out of that, and attempting to will cost you attention you need for the task. Perform the station to the published criteria and move on.

3. Preparing unevenly

Because reliability comes from sampling across stations, [1][3] the exam rewards a reliable floor rather than a high ceiling. Two competent stations beat one excellent and one poor on almost any marking scheme. Identify your weakest station type and spend your time there, not on polishing what you already do well.

4. Treating communication as the easy marks

Interpersonal skills were assessed less reliably across stations than clinical skills, [1] which means your communication mark will swing more between stations than your technique mark. It also means communication cannot be crammed as a set of phrases — it is being judged as a general trait that should hold across very different situations.

5. Assuming a high score means high competence

Overall reliability of 0.66 across stations [1] means a meaningful share of your rank is noise. That cuts both ways: a good OSCE score is encouraging but is not proof you are safe, and a mediocre one is not proof you are not. Use placement feedback, which samples you over weeks, as the better signal.

6. Thinking the format is unfair rather than imprecise

These are not the same complaint. The OSCE is a deliberately structured attempt to reduce bias, and well-designed versions reach G-coefficients above 0.8. [4] Its weakness is precision from limited sampling, which schools address by adding stations — a minimum of 17 for defensible pass-fail decisions in one analysis. [3] Knowing the mechanism is more useful than resenting it.

What the evidence supports — and what it does not

Supported

  • Preparing for consistency across station types rather than excellence in one. [1][3]
  • Treating a single poor station as expected variance. [3]
  • Practising communication in varied scenarios, since it does not generalise well. [1]
  • Working to the published marking criteria rather than to a perceived examiner preference. [2]
  • Trusting longer sampling — placement feedback — over a single exam. [1]

Not supported

  • Reading an OSCE score as a precise measure of competence. Alpha 0.66. [1]
  • Assuming the examiner is a neutral instrument. Examiner variance exceeded student variance fourfold. [2]
  • Short OSCEs for high-stakes pass-fail decisions. 17 stations were needed in one analysis. [3]
  • Treating communication marks as more stable than technique marks. The reverse held. [1]

How certain is this?

Evidence grade: Moderate.

The core claim — that OSCE reliability is driven by sampling and that examiner and station effects are large relative to candidate ability — is supported by a meta-analysis of 188 coefficients from 39 studies [1] and by two independent variance component analyses using different statistical frameworks. [2][3] Agreement across methods is what makes it credible.

What limits the grade is transfer. None of these studies is in physiotherapy: they cover medical, dental and nursing examinations. [2][3][5] The mechanism is generic to the format rather than to the profession, but the specific percentages and station counts should not be quoted as physiotherapy figures.

A note on the dates. Three of these papers predate 2021. Generalisability analyses of an assessment format do not date the way treatment evidence does, and the most recent study here [4] reaches the same conclusions as the oldest.

Common questions

How should I actually prepare?

Broadly and evenly. Reliability comes from sampling across stations, [1][3] so your expected score depends more on your weakest station type than on your strongest. List the station types your programme uses, rank yourself honestly, and work from the bottom up. Rehearse out loud with a partner marking against the published criteria, not silently.

I got a harsh examiner. Does that really matter?

Measurably, yes — examiner stringency variance was more than four times examinee variance, and adjusting for it changed pass-fail status for about 11% of candidates in one study. [2] That is a reason for schools to use two examiners per station, which the meta-analysis found improves reliability. [1] It is not something you can prepare around, so do not try.

Is it worth appealing a single bad station?

Follow your institution's process, but understand the statistics: station difficulty and student-station interaction together dominated the variance in one large analysis. [3] Examiners and schools know a single station is a weak signal, which is precisely why decisions are made across the whole exam rather than station by station.

Why does my communication mark jump around?

Because it genuinely is less stable. Interpersonal skills were evaluated less reliably across stations than clinical skills, [1] which reflects that communication is situation-dependent in a way that a technique is not. Practise it across deliberately different scenarios — anxious patient, rushed handover, family present — rather than rehearsing one script.

Does this apply to the viva too?

The same principles apply more sharply, because a viva samples less. Fewer examiners and fewer questions mean examiner effects have more room to act, and the OSCE literature's remedy — wider sampling [3] — is exactly what a viva lacks. Prepare to reason aloud in a structured way, and expect variability in the questions you happen to be asked.

References

  1. Brannick MT, Erol-Korkmaz HT, Prewett M. A systematic review of the reliability of objective structured clinical examination scores. Medical Education. 2011 Dec;45(12):1181–9. doi:10.1111/j.1365-2923.2011.04075.x PMID 21988659 Meta-analysis of reliability coefficients
  2. Harasym PH, Woloschuk W, Cunning L. Undesired variance due to examiner stringency/leniency effect in communication skill scores assessed in OSCEs. Advances in Health Sciences Education. 2008 Dec;13(5):617–32. doi:10.1007/s10459-007-9068-0 PMID 17610034 Variance components analysis
  3. Schoonheim-Klein M, Muijtjens A, Habets L, et al. On the reliability of a dental OSCE, using SEM: effect of different days. European Journal of Dental Education. 2008 Aug;12(3):131–7. doi:10.1111/j.1600-0579.2008.00507.x PMID 18666893 Generalisability study
  4. Hara S, Ohta K, Aono D, et al. Feasibility and reliability of the pandemic-adapted online-onsite hybrid graduation OSCE in Japan. Advances in Health Sciences Education. 2024 Jul;29(3):949–965. doi:10.1007/s10459-023-10290-3 PMID 37851159 Generalisability study
  5. Castro-Yuste C, García-Cabanillas MJ, Rodríguez-Cornejo MJ, et al. A Student Assessment Tool for Standardized Patient Simulations (SAT-SPS): Psychometric analysis. Nurse Education Today. 2018 May;64:79–84. doi:10.1016/j.nedt.2018.02.005 PMID 29459196 Instrument validation study

About this resource

How to use this

Written to be learned from, not memorised.

This page tells students where their OSCE score actually comes from, so their preparation goes to the part they can change. Faculty may use this page in teaching with attribution. It carries its review date and its next review date, so you can see at a glance whether it is current before you put it in front of a cohort.