AI in assessing people: where the usefulness stops
Echipa HR 365 · reviewed 2026-09-10 · 5 min read
The test that settles everything: can you explain to the person why they got what they got, in their own terms, with examples? If the answer depends on a score you cannot decompose, you do not have a review — you have a number you are communicating. The difference is immediately obvious to the person being reviewed.
Why a generated score does not work
Three practical reasons, independent of the quality of the model. First: a review has a communication function, not only a measurement one — and what cannot be explained communicates nothing. Second: the data it would rest on are usually poor proxies — message counts, hours online, tasks closed. Third: from the moment people know what is being measured, behaviour shifts towards the indicator and detaches from its purpose.
What does help
| The task | What it does | Why it is safe |
|---|---|---|
| Summarising a half-year of one-to-one notes | Gathers what you wrote, by theme | The raw material is your observation, not an inference |
| Suggesting questions for a conversation | Proposes, you choose | It produces no assessment |
| Rewording feedback that came out too harsh | Adjusts the tone, keeps the fact | The fact stays yours |
| Flagging an unfilled criterion | A formal check | It does not touch the content |
| Grouping survey answers | Recurring themes, at an aggregate level | It is not applied to a person |
The common denominator: in all of them, the raw material is something a person observed, and the output is help with drafting or organising. From the moment the model starts inferring a quality of the person from indirect data, the line has been crossed.
The line, stated plainly
- Organising what a person observed: acceptable.
- Rewording what a person decided: acceptable.
- Inferring a quality from indirect data: no.
- Producing a score that influences a decision about the person: no.
Point three covers most commercial products promising “performance analytics”. What they actually do is infer engagement from observable activity — and observable activity is the most easily imitated part of the work.
The case of surveys and free text
Here AI is genuinely useful: grouping two hundred free-text answers into recurring themes is work nobody does well by hand. The condition is that the result stays aggregated and that individual answers are attributed to nobody.
What is never done: tone or sentiment analysis on named responses, or on internal messages. That is surveillance, and the label does not change its nature — and once discovered, it destroys trust in every future survey.
What you tell people
Three sentences, said once and honoured
- What is used — where exactly AI helps in HR processes — concretely, not in general
- What is not done — no scores are generated about people, and no messages or tone are analysed
- Who decides — every decision about a person is made by a person who owns it
The second sentence matters most, because it is exactly what people assume is happening when they hear “AI in HR”. Left unsaid, the assumption stands — and it affects even how honestly they answer surveys.
But what if the model is more objective than the manager?
It may be more consistent, which is not the same as more correct: consistently wrong is still wrong, only systematically. And the remedy for inconsistent managers is written criteria and independent rating, not replacing judgement.
Can I use AI to prepare the review conversation?
Yes, and it is one of the best uses: give it your notes from the whole half-year and ask for a structure. What comes out is the order of the conversation, not its content.
Where to start
Check whether any tool you already use produces a score about people. There usually is one, in an adjacent platform, and nobody has looked at what it calculates.
Then tell the team the three sentences. It takes five minutes and removes an assumption that otherwise contaminates all the data you collect from them.
What to ask a vendor promising performance analytics
- From exactly what data is it calculated? If the answer includes observable activity — messages, hours, clicks — you already know what it measures.
- Can the product explain an individual result, with examples, to the person being assessed?
- What happens if someone disputes a score? Is there a route to verification, or only a recalculation?
- Can the individual-score part be switched off entirely, keeping the rest?
The fourth question is the most practical: many products bundle genuinely useful functions — summaries, organisation, aggregation — together with individual scores. If they can be separated, the discussion becomes far simpler.
What changes when people find out
The moment a team learns that a tool produces scores about them, two things change immediately: honesty in surveys and the behaviour being measured. Both degrade, and recovering trust takes far longer than installing the tool.
Which is why, if something of this kind is introduced anyway, the only viable route is to announce it beforehand, explicitly, with what is measured and how it is used. Discovered afterwards, it produces a far bigger problem than any efficiency gain.
Conversation summaries and suggestions, with no scores about people
AI organises what you noted and proposes the shape of the conversation; the assessment stays the manager’s, with the criteria and the evidence in plain sight.
You can create an account in a few minutes and use every module for 7 days, no card required.