If you are responsible for assessing hands-on skills – in a fire academy, an allied health program, an ambulance service, a manufacturing plant – you have almost certainly come to recognize, and even (perish the thought) accept, the fact that different assessors score differently, according to their own experiences, biases and preferences. But we should not accept this. The cost is too great in terms of safety, efficiency and reputation. Fortunately, measurement science has been thinking about it for decades. The solution comes down to two terms: “validity” and “reliability.” When creating a skill assessment, it is very useful to have these terms in mind, to understand how one of them depends on the other, and to see why the second one – reliability – is both the more neglected and the more fixable of the two.

Long before founding SkillGrader, I spent a decade as a Computer Science faculty member at UBC, where my research area was online learning and assessment. Validity and reliability were the bedrock of all assessment experimentation and implementation, and they are the bedrock of the work we are doing at SkillGrader. Let’s make them a part of your assessment practice if they are not already.

What do these terms mean?

Validity asks a simple question: does the assessment measure what it claims to measure? If your skill sheet for patient assessment actually captures whether a student can assess a patient – rather than, say, whether they can recite steps in a memorized order – it is valid. Most organizations intuitively care about validity, and they express that care through the content of their assessment forms. When we debate which steps belong on a skill sheet, and which are critical failures, we are doing validity work whether we know it or not.

Reliability asks a different question: would the same performance receive the same score regardless of who assessed it, or when? If Assessor A passes a candidate that Assessor B would have failed, the assessment is unreliable. The score reflects the assessor as much as the performance – and it should not.

There is the relationship between validity and reliability: an unreliable assessment cannot be valid. If the same performance produces different scores depending on who happens to be holding the clipboard, then the score cannot be measuring the performance. It is measuring something else – the assessor’s mood, their standards, their interpretation of an ambiguous rubric. You can have a perfectly designed, expert-vetted skill sheet, and if two assessors apply it differently, the results mean far less than everyone believes they do.

Why does reliability get so little attention?

The answer, I believe, is that on paper, reliability can be invisible. Validity problems, by contrast, are usually pretty clear – someone eventually notices that the skill sheet is missing something important, and it gets updated. Domain experts (like you) are generally very good at validity. You know your topic, and you know what trainees or students need to do to demonstrate proficiency.

Reliability problems hide. When assessments live on paper in filing cabinets, there is simply no practical way to ask the question “do my assessors agree with one another?” The data are there, in principle. In practice, nobody is going to hand-tabulate thousands of skill sheets to compute agreement between assessors. Being good at reliability also requires assessment science expertise, which most people do not have – not skill domain expertise, which you have. So the question never gets asked, and an organization can run for decades on assessment results whose consistency has never once been examined and is not understood.

What can we do about it?

The good news is that reliability responds well to structure and measurement. Well-defined observable indicators – rather than broad judgment calls – narrow the room for interpretation. Two assessors are more likely to agree on clear, observable facts than on whether something constituted a 3 out of 5 or a 4 out of 5. So it is very much about asking the right kind of question: one where the answer is hard to get wrong when observing a trainee.

And most importantly, digital assessment makes reliability measurable for the first time. When every assessment is captured as structured data, you can finally see whether one assessor scores systematically harder than the others, or whether a particular indicator produces scattered results because it is ambiguously worded. This is precisely the kind of visibility we built SkillGrader to provide. What you can see, you can fix – through assessor calibration, rubric revision, or both.

The bottom line is this: your validity work is only as good as your reliability, and your reliability is only as good as your ability to structure your assessment correctly and measure the outcome. The tools to help with both of these now exist.

Thanks for reading. Until next time.

About the author

Murray Goldberg is the founder and CEO of SkillGrader, a platform for objective observational skill assessment. A former tenured faculty member in Computer Science at the University of British Columbia, Murray's research area was learning technologies, and in 1995 he created WebCT — the first widely-used learning management system in higher education, eventually serving 14 million students in 80 countries. He has spent three decades working to advance the art and science of learning and assessment.