The review form: the scale, the criteria and why the overall average is useless
Echipa HR 365 · reviewed 2026-09-02 · 5 min read
A useful review form has five or six criteria tied to the real work, a scale with written descriptions for each level, and no aggregate final score. An average across criteria looks objective and is precisely the mechanism by which a review stops meaning anything: it hides the very criterion the conversation should be about.
The criteria: from their work, not from a catalogue
The most frequent defect of review forms is that they are the same across the whole company: “communication”, “teamwork”, “results orientation”. They sound right and cannot be rated, because nobody knows what “communication 4” looks like for an accountant compared with a field agent.
Useful criteria are written starting from what the person does in an ordinary week. The practical rule: if you cannot give an observable example from the last three months for a criterion, it does not belong on the form.
| Generic | Ratable | What you observe |
|---|---|---|
| Communication | Flags in advance when a piece of work will be late | How many delays were flagged before the deadline |
| Results orientation | Carries work through to the recipient’s confirmation | How many stopped at “I sent it, I am waiting” |
| Teamwork | Picks up tasks when a colleague is away, without being asked | Concrete situations from the period |
| Attention to detail | Work does not come back for formatting corrections | The return rate, if it exists; otherwise examples |
| Initiative | Proposes changes to the things they do daily | What they proposed and what happened to the proposal |
The scale: four levels, each described
Five levels produce a comfortable middle where half the organisation lands. Four levels force a position. What matters is not the number but the description: without it, “3” means something different to every manager, and comparisons between teams become fiction.
| Level | Description | What the assessor must write |
|---|---|---|
| 1 — below requirement | The result needs somebody else’s intervention to be acceptable | Two concrete situations and what happened |
| 2 — nearly at requirement | The result is good with supervision or minor corrections | What still needs attention |
| 3 — at requirement | Does what the role involves, autonomously and consistently | Nothing required; it is the expected state |
| 4 — above requirement | Also resolves things outside the role, and they stay resolved | Two concrete situations, with a visible effect |
Why you do not calculate the average
Somebody with 4, 4, 4, 4 and 1 averages 3.4 — the same as somebody with 3, 3, 4, 3 and 4. The first situation calls for an urgent conversation about the criterion where they scored 1; the second is solid performance with nothing urgent. The average makes them look identical and moves the discussion from “what do we do about this” to “what mark did I get”.
What replaces the average: two or three sentences of conclusion written by the assessor, saying what to keep, what to change and what support the person gets. The conclusion is written after the ratings, not before — otherwise the ratings get adjusted to support it.
The process, in the order that matters
A review cycle that does not become a formality
- Employee — completes a self-assessment against the same criteria, before seeing the manager’s
- Manager — completes theirs independently, without reading the self-assessment, so as not to anchor
- Manager and employee — compare the two forms; large differences on a criterion are the subject of the conversation, not the average
- Manager — writes the conclusion and one concrete action for the period ahead
- HR — checks only whether the extreme levels have the required evidence, not the content of the review
The order of the first two steps is the one non-transferable part. A manager who reads the self-assessment first rates, almost systematically, closer to it than they would have alone — and then the exercise becomes a negotiation rather than a review.
Three predictable rating biases
They appear in almost every assessor and have nothing to do with bad faith. Once named, they are easy to correct:
| Bias | What it looks like | The correction |
|---|---|---|
| Recency | The last few weeks count, not the whole period | Ratings are written from one-to-one notes, not from memory |
| The halo effect | One strong quality lifts every other criterion | Criteria are rated one at a time, across everybody, not person by person |
| The comfortable middle | Almost everybody gets the middle level | An even-numbered scale and required evidence for the extremes |
The second correction is the most effective and the most rarely applied. By rating one criterion across everybody at once, the assessor compares comparable behaviours and notices the differences themselves — instead of building, person by person, a general impression which then gets distributed across criteria.
What happens after the review
A review with no follow-through is the surest way of teaching people the exercise is a formality. The written conclusion has to yield exactly one thing the person will do differently and one thing the manager will do differently — both verifiable at the next conversation.
The sizing rule: one action each. Three development objectives set simultaneously means, in practice, zero, because none has priority. And the manager’s action written alongside the employee’s changes the nature of the conversation: it is no longer a verdict, it is a mutual commitment.
How often should reviews happen?
Twice a year is a reasonable compromise: rare enough for there to be material, frequent enough for the conversation to still change something. Annual works only if monthly one-to-ones handle the ongoing corrections.
Do I link the review to pay?
If you link it directly and automatically, the review becomes a negotiation and people stop acknowledging what is not working. Most companies that work well keep a link, but with a separate conversation, offset in time.
Review cycles with a separate self-assessment and a history per person
The self-assessment and the manager’s review are completed independently, and the comparison appears automatically — so the conversation starts from the differences, not from an average.
Free account, every module for 7 days, no card required.
Sources
- Daniel Kahneman, Olivier Sibony, Cass Sunstein — Noise — the chapters on performance evaluation: variability in ratings between assessors falls when criteria are described explicitly and rated independently, before the joint discussion. · read on 2026-09-02