Work
Letting an Agent Draft Performance Reviews: The Line Worth Drawing
Where an agent drafting performance reviews helps, where it quietly replaces the manager's judgement, and how to tell which side of the line you are on.
Written by Sicherhaven
Review season arrives and every manager has eight reviews to write and no memory of March. So someone suggests an agent draft them. The instinct is reasonable, and the risk is specific.
Here is the line worth drawing. An agent drafting performance reviews may gather evidence and tidy phrasing. It should not decide what the evidence means. The moment the model is the thing forming an opinion about a person's year, you have swapped a manager's judgement for a plausible sentence, and nobody in the room will be able to tell the difference by reading it.
The two halves of writing a review
Writing a review is really two jobs stuck together.
The first job is retrieval. What did this person actually work on? Which projects did they carry, what shipped, what slipped, what did they pick up that was not theirs. This job is tedious, it depends on records nobody remembers to keep, and managers are bad at it because human memory over twelve months is bad.
The second job is judgement. Given all that, how did this person do? What should they do differently? What are you willing to say to their face and stand behind in a promotion committee.
An agent is good at the first job and has no business doing the second. The same split runs through hiring, where some steps are safe to delegate and some are not.
What is fair to hand over
These uses keep the manager's judgement intact:
- Pulling the year's work into one place: tickets closed, projects owned, dates, who they worked with.
- Flagging things the manager may have forgotten, especially work from the first quarter.
- Turning the manager's rough notes into clean prose without changing the meaning.
- Checking that the written review does not contradict itself, and that every claim of a problem has a concrete example attached.
- Making the tone consistent across eight reviews written at eleven at night in different moods.
Each of those takes something the manager already decided and makes it clearer. None of them decide anything.
Where it stops being help
The problems start when the prompt sounds like "assess this person's performance based on their record".
A model given a year of activity data will produce an assessment. It will read fluently. It will use the vocabulary of your review template. And it will be built on whatever is visible in the records, which is not the same as whatever mattered.
Records over reward the visible. The person who closed forty small tickets looks busier than the person who spent three months on the one problem that was blocking everyone. The person who works in shared documents leaves more traces than the person who does their thinking in conversation. A manager knows this. A model reading the record does not.
There is also the reverse failure. A manager who reads a generated assessment before forming their own tends to anchor on it. They spend the hour editing the draft's opinion instead of arriving at theirs. The review then looks like the manager's work and is not.
A simple test
Before letting an agent touch a review, ask one question about each thing it produces: if this turns out to be wrong, would the manager have caught it?
Evidence gathering passes. If the agent misses a project, the manager notices the gap because they were there.
Phrasing passes. If the sentence reads oddly, the manager sees it.
An overall rating fails. If the agent's judgement is subtly off, there is nothing for the manager to check it against except a judgement they have not formed yet.
Sequence matters more than rules
The practical fix is ordering. Have the agent produce the evidence pack first, with no assessment in it. The manager reads the pack and writes their own view in whatever rough form suits them. Only then does the agent help with wording.
That order keeps the model working on the parts where being wrong is visible, and keeps the manager doing the part where being wrong is not.
What this looks like in a system
SicherOne puts project management and HR on one set of records, which is why the evidence gathering half is workable at all: the agent has the project history and the leave record in the same place, so it can tell the difference between a quiet quarter and a quarter someone spent on parental leave. A human approves agent output before it ships, which is the enforcement point for everything above. The manager is not reviewing a draft they can rubber stamp. They are the author.
It is one reason HR decisions sit near the bottom when departments are ranked by risk. Whether an agent may draft a review at all is also a matter of local employment law and your own policy, and rules on automated decision making about employees differ by country. Check with your legal advisers before wiring anything into a formal process.
The uncomfortable part
Some managers want the agent to write the whole thing because writing reviews is the part of management they dislike most. That is honest and it is worth naming. Picking the most disliked task first is a familiar pattern, and not always the one that pays off.
The response is not a better tool. It is that forming a view on someone's year, in words you are willing to defend, is a large part of what a manager is for. Handing it to a model does not save the work. It just moves the moment when someone notices the work was never done.
← All postsWe're building the future of community events and financial wellness
See how Eventify and WealthWise change the way people find events and manage money.
Get Started
