Work
Setting Confidence Thresholds Without Reading a Research Paper
A manager friendly way to set confidence thresholds for when an agent escalates: start cautious, watch your rejection rate, and move the line from evidence.
Written by Sicherhaven
Somebody asks what confidence level the agent should need before it acts on its own instead of asking a person. Nobody in the room knows, so a number gets picked because it sounds sensible, and it stays there for a year.
Setting confidence thresholds does not require a research paper. Start deliberately too cautious, so almost everything escalates, then watch the rejection rate on what humans approve and move the line only where the evidence says the agent was reliably right. The number is an output of watching, not an input you guess.
Why picking a number fails
Two reasons, and both are practical rather than technical.
First, a confidence score is not a promise. It is the model's own sense of how sure it is, and models can be sure and wrong. The relationship between the score and actual accuracy differs by task, by data and by how the question is phrased. So a number that works well on one workflow tells you very little about another.
Second, nobody in the meeting has any basis for the choice. The number that gets picked reflects how nervous the room feels that day. Then it becomes settled policy because changing it would require the same conversation again.
Start where escalation is annoying
Set the threshold high enough that nearly everything goes to a human. Yes, this is inefficient. It is inefficient on purpose and only for a few weeks.
What you get in exchange is a body of evidence: for every case the agent handled, a human decision about whether it was right. That is the only data that answers the question you actually have, and it accumulates on its own once the work is split into a draft step, a review step and a send step.
Two things to record for every escalated case:
- The agent's confidence score.
- Whether the human accepted the output unchanged, changed it, or rejected it.
That is enough. No statistics training required.
Read the rejection rate by band
After a few weeks, group the cases into bands by confidence and look at the acceptance rate in each.
You are looking for a level above which humans almost never change anything. That is your candidate threshold. Below it, escalation is earning its cost, assuming the approval step is written so people actually read it. Above it, the human is a formality.
The shape people usually find is that the very high confidence band is genuinely reliable, the middle is mixed, and the low band is worse than the score suggests. Where the bands sit is specific to your work, which is precisely why guessing does not work.
If there is no band with a clean acceptance rate, that is a real answer too. It means this task is not one to automate past a human yet, and the honest move is to leave the threshold where it is and revisit later.
Move the line slowly, and only down
When you lower the threshold, lower it one band, and keep watching the same numbers. If the rejection rate in the newly automated band rises above what you accepted before, put it back.
Some rules that keep this sane:
- Change one threshold at a time, so you can tell what caused what.
- Wait long enough to see a normal cycle of work before judging.
- Never raise automation on the back of a quiet period.
- Set a review date at the start, so the number does not become permanent by default.
Carve out the cases that always escalate
Confidence should not be the only gate. Some cases go to a person regardless of how sure the agent is.
Anything with money in it. Anything about a specific employee. Anything going outside the company. Anything where the agent found a conflict between two sources. Anything unusually large compared to the normal case.
These are category rules and they sit above the threshold. A high score on a risky category means the agent is confident about something you were never going to let it do alone.
Watch for the quiet failure
The number that should worry you is a rejection rate that falls to nearly zero across every band at once.
Sometimes that means the agent improved. More often it means approval has slid into rubber stamping, because the last hundred were fine and the queue is long. The threshold data is only meaningful while the human decisions behind it are real.
A crude check: sample a handful of approved cases each month and have someone else review them properly. If that second review finds problems the first missed, your threshold data has been measuring attention, not accuracy.
Where it lives
SicherOne has a human approving agent output before it ships, and agent activity sits on the same records as the project and HR side, so the acceptance and rejection history you need for this exercise accumulates as a side effect of normal work rather than as a separate logging project.
The habit to keep is the quarterly look. Thresholds set once and never revisited drift out of date the same way any other configuration does, and the drift is invisible until the day it is not.
← All postsWe're building the future of community events and financial wellness
See how Eventify and WealthWise change the way people find events and manage money.
Get Started
