The Skill Lab is a series of simple, science-backed tips for anyone who trains or tests practical skill. If you build skill sheets, watch people perform, or decide who passes, it is written for you.

How many of us give much thought to what constitutes a valid “indicator” or observation on a skill assessment sheet. But we should. If you pick up a skill sheet from your own program and read down the list of indicators - what do they look like? Some of them may describe something a person visibly does. Others may describe a conclusion you would have to draw about them. On paper the two look similar but they behave very differently the moment two assessors take that sheet out to assess a trainee. One is objective and one is subjective. One kind of indicator results in consistent data gathering and assessment. The other produces inconsistency - meaning we can’t rely on the results of our assessments. If we cannot rely on the results, we have no way to tell if our training is working.

The “correctness” of an indicator decides whether a score tells you something about the trainee or something about which assessor happened to be working that day. And when the score is the basis of a pass, a promotion, or a decision to put someone on the job, that distinction is important.

What is an indicator actually for?

An indicator has one job: to tell us objectively whether something happened. It is not there to tell us how good that something was. Judging quality is the work of the layer above - the rubric, the weighting, and the scoring rules that decide what a particular pattern of recorded actions adds up to.

Said another way, a well-built indicator is a granular record of whether a specific action occurred. Was the step performed? Was the phrase spoken? Was the check made before the patient was moved? A form built from indicators like those produces something a form built from verdicts won’t - a detailed factual record of what actually took place. That record is what can then be analysed, consistently and objectively, to answer the question we actually care about: what was the quality of this performance?

An easy test is to look at your indicators and simply ask yourself whether two assessors would need to confer in order to agree on the result? Two people are very likely to agree on whether a sentence was spoken or a step was performed in the right order. They are much less likely to agree on whether what they saw added up to “good situational awareness” - not because either of them is careless or incompetent, but because each is comparing the performance against their own mental picture built over a different career.

Why our sheets fill up with conclusions

There is a reason judgement-based indicators end up on skill sheets. “Demonstrates professionalism” is on the sheet because professionalism genuinely matters, and because the person who wrote that line is an expert who can recognize it instantly and reliably in the field. The line is a compression of real expert judgment, written by someone who has that judgment and likely assumed the next assessor would too.

The problem is not the expertise. It is that the judgment gets made privately, by a different person each time, using criteria nobody has written down - or at best is listed in a rubric on the back of the sheet that few people refer to or interpret consistently. Separating the record of what happened from the rules that decide what level of quality it demonstrates is assessment-science work - a different skill, and one most of us had no particular reason to learn.

What to do about it

The conversion is usually straightforward once you see the pattern. “Demonstrates professionalism” becomes something closer to: introduces themselves to the patient by name and role; explains what is about to happen before touching the patient; addresses crew members by name. “Manages the scene effectively” becomes: assigns a task to each crew member before entry; states the operational objective aloud. Each of those is a fact rather than a verdict.

A reasonable objection at this point is that we have thrown away the expert judgment that made “demonstrates professionalism” worth writing in the first place. We have not. We have simply moved it to the scoring rules, where it can be written down once, applied identically to every candidate, revised when we learn something, and explained to a trainee who asks why they did not pass. That is a considerably better home for it than the private impression of whichever assessor performed the assessment.

And the bonus good news is that trainees are more likely to appreciate and trust assessments that are based on objective indicators, and where the score is generated programmatically rather than subjectively. This results in a more collaborative and positive debrief. We have actually done surveys on this, and the results are really positive. More on that in some future edition of The Skill Lab!

Until next time, thanks for reading and keep well.

About the author

Murray Goldberg was a tenured faculty member in Computer Science at the University of British Columbia, where his research into learning technologies led him to create WebCT in 1995 — the first widely-used learning management system in higher education, eventually serving 14 million students in 80 countries. He has spent three decades working to advance the science and practice of learning and assessment, and is the founder and CEO of SkillGrader.