Skip to main content

Industry

Counting the Minutes: Measuring Time Saved in a Way That Survives Scrutiny

A sampling method for before and after task timing that a finance team will accept, including how to subtract review time without flattering the result.

Written by Sicherhaven

Someone in your team says the new tool saves about two hours a week per person. Finance asks how you know. The answer is that people said so, and the conversation ends there.

Measuring time saved in a way that survives scrutiny takes a sample of real tasks timed before the change and the same tasks timed after, with review time subtracted from the after figure. It takes a couple of weeks and produces a range rather than a number. That is less impressive and far more likely to be believed.

Sample, do not survey

Asking people how much time they save produces an estimate shaped by how they feel about the tool. Enthusiasts overstate, sceptics understate, and both are answering honestly.

Instead, pick one task type and time actual instances of it. Twenty is plenty for most teams. Ten is enough to see whether the effect is large or small, which is often all you need to decide.

Time it with a start and a stop, recorded by the person doing it. Not a stopwatch running all day, not calendar inference, not a guess at the end of the week. A note that says started 10:12, finished 10:41.

Do this before you change anything, alongside the other numbers worth capturing ahead of a rollout. A before measurement taken after the change, from memory, is the single most common way these exercises get dismissed.

Measure the same task, not the same feeling

Define the task narrowly enough that two instances are comparable. "Writing a client update" is too loose. "Writing the fortnightly update for an active project, from status notes to sent" is measurable.

Also record something about size, because task length varies more than tool effect. A word count, a number of line items, a number of people involved. If your after sample happens to contain smaller tasks, the size field is what saves you from claiming a saving that was really luck.

Keep the same people in both samples where you can. Different people work at different speeds, and a change of staff between measurements will swamp the thing you are trying to see.

Subtract review time honestly

This is where most internal measurements fall apart, and it is the first thing a finance team will ask about.

When an agent drafts something and a person reviews it, the after time is the draft time plus the review time plus any rework. Not the draft time alone. A workflow where the machine produces something in seconds and a person spends twenty minutes fixing it has not saved twenty five minutes.

Be specific about what counts as review:

  • Reading the output.
  • Checking facts against a source.
  • Editing.
  • The second pass after an edit, if there is one.
  • The approval click itself, which is small but real.

Also count the times output was rejected and the task was done from scratch. Those instances belong in the after sample. Dropping them because they are not representative is exactly what makes a number unbelievable.

What finance will ask

Prepare for four questions, because they come up nearly every time.

How many instances did you measure, and over what period. Small samples are fine if you say they are small.

Who did the measuring, and did they have an interest in the result. A person championing the tool timing their own tasks is not disqualifying, but it should be disclosed.

Does the saved time land anywhere. Twenty minutes saved across forty people is not a headcount. Counting hours without asking where they went is one of the signs a pilot is measuring the wrong thing. It is capacity, which is real but different, and calling it a cost saving invites the response that nobody's salary went down.

What is the ongoing cost. Licence, setup, training and the review time you just counted. A saving reported without the cost beside it reads as advocacy.

Report a range, and report what failed

Give a range rather than a point. "Between eight and eighteen minutes per update, from a sample of twenty two" is a sentence that holds up. "Saves 34 percent" is a sentence that invites someone to ask where the number came from, and there will not be a good answer.

Include the cases where it did not help. Every honest measurement has some. Reporting them makes the rest of the number credible, and it usually points at the real finding, which is that the tool helps a lot on one kind of task and not at all on another.

The version that takes a week

If two weeks of measurement is more than the decision deserves, do the short version, or fold the timing into a two week trial built to produce a real yes or no. Pick one task type, time five instances by hand before, five after, subtract review time, report the range and say plainly that the sample is five.

That is a weaker claim than a proper study. It is a much stronger claim than a feeling, and it is honest about which one it is. Most internal decisions do not need more than that.

← All posts

We're building the future of community events and financial wellness

See how Eventify and WealthWise change the way people find events and manage money.

Get Started