The end-of-course test: writing questions that measure something
Echipa HR 365 · reviewed 2026-09-06 · 5 min read
A good test question asks for application, not reproduction. The practical difference: “What are the three stages of the process?” measures whether the person read the last slide; “A client calls and says X. What is your first step?” measures whether they can use what they learned.
What you actually want to measure
An internal test has one useful purpose: to find out whether the person can do the work after the course. Not whether they paid attention, not whether they went through the material, not whether they can reproduce definitions. All three of those are measured more cheaply another way — through course progress, which is visible anyway.
The consequence: if a question does not describe a situation the person could find themselves in, it probably measures nothing useful. That is the test worth applying to every question you write.
Six patterns that work
| Pattern | Example wording | Good for |
|---|---|---|
| Situation and first step | “X happens. What do you do first?” | Procedures, order of operations |
| Choosing between close options | “Which of these two approaches fits, and why?” | Judgement, not memorisation |
| Spotting the error | “What is wrong in this situation?” | Checking, attention to detail |
| The edge case | “What do you do if Y is missing?” | Real understanding, not a learned pattern |
| Ordering | “Put the steps in the correct order” | Processes with a required sequence |
| Applying a rule to data | “With these figures, what is the result?” | Calculations, thresholds, criteria |
What does not work
- Questions with “all of the above”. They are guessable and measure nothing.
- Obviously false distractors. If three of four options are absurd, the question is decorative.
- Negative phrasing: “Which is NOT…”. It tests reading attention, not knowledge.
- Questions about exact figures the person can look up any time. They measure memory, not competence.
- Questions that depend on a particular phrasing from the course. Whoever understood it differently but correctly fails.
The last point is the subtlest and the most unfair. It appears when the question is written by copying a sentence from the material: the correct answer is the one that most resembles the text, not the one that is true. The test then becomes an exercise in recognising the author’s style.
How many questions and what pass mark
Three questions at the end of each module work better than twenty at the end: they catch the misunderstanding where it happened and do not turn the end of the course into an exam. A final test, if there is one, checks the connections between modules rather than repeating them.
What a pass mark means
A 10-question test with 4 options each, one correct. Somebody guessing entirely.
- Probability of hitting one question
- 25%
- Expected score from pure guessing
- 2.5 out of 10
- A 60% pass mark
- hard to reach by guessing, easy with a superficial read
- An 80% pass mark
- requires real understanding; also produces retakes
- The reasonable mark for a mandatory course
- 70–80%, with retakes allowed
Nu intră în calcul:
- questions a competent person can get wrong because of the wording — those get rewritten, not compensated for with the pass mark
Allowing retakes matters more than the pass mark. A test that can be retaken after re-reading the module turns the check into learning; one with a single attempt produces anxiety and answers copied between colleagues.
Where the result is computed
A technical detail with practical consequences: if passing is decided in the learner’s own page, the result is a claim they make, not a measurement. For an internal course with no stakes, it does not matter. For one that produces a certificate or an authorisation to do something, it matters a great deal.
Where to start
Take the test from an existing course and count how many questions describe a situation. If it is under half, you already have your rewrite list.
Then look at the questions most people fail. Usually one or two dominate, and usually the problem is the wording rather than the material.
What you do with the results
The results of an internal test say two different things and it is easy to confuse them: something about the learners and something about the course. The second is usually more valuable and almost never used.
| What you observe | What it means | What you do |
|---|---|---|
| A question most people fail | Either it is badly worded, or the module does not explain it | You re-read both; usually it is the wording |
| A question nobody gets wrong | Too easy, or the answer is obvious from the wording | You rewrite it or drop it |
| High scores everywhere | The test checks short-term memory | You add edge-case questions |
| A big gap between teams | The application context differs | You check whether the course is relevant to both |
The first row is worth checking every time. In a ten-question test, one question usually dominates all the failures — and in most cases the problem is an ambiguous wording, not a gap in understanding.
How you organise retakes
An immediate retake with the same questions measures ten-second memory. What works: a retake allowed after re-reading the module, with questions from the same bank but not identical. That requires a question bank larger than the test, which is effort at first construction and zero at every later retake.
Do I test on non-mandatory courses too?
Yes, but with no pass mark. Its role there is not to filter but to give the learner a signal about what they understood — and you a signal about what the course does not explain well.
How do I stop answers being shared between colleagues?
A question bank larger than the test, and shuffled ordering. It does not eliminate the phenomenon, but it makes it more expensive than going through the module.
Question banks, with results computed on the server
Questions are reused across courses, and the correct-answer rate on each shows you which are badly worded — not just who failed.
Free account, every module for 7 days, no card required.