Skip to main content

Work

How to Run a Two Week AI Trial That Produces a Real Yes or No

A two week AI trial design with one named decision maker, one workflow, a success number agreed in advance, and a written rule for stopping the pilot early.

Written by Sicherhaven

Most AI trials end with a meeting where everyone says it was interesting and nobody says yes or no. That happens because the trial had no decision maker, no single workflow and no number agreed in advance. A two week AI trial that produces a real answer needs four things fixed before the software is switched on: one named person who decides, one workflow, one success number, and a written rule for stopping early.

Set those four before day one and the trial ends in a decision. Skip any of them and it ends in another trial.

Name one person who decides

Not a committee. One person, named in writing, who will say yes or no on the final day and whose judgement everyone has agreed to accept beforehand.

This person should be close to the work. A team lead who runs the workflow every week will spot a bad output that a director would wave through. If the decision needs sign off higher up, that is fine, but the recommendation comes from one named person, not from a group discussion.

Pick one workflow, not a tool

The mistake is trialling a product. You cannot evaluate a product in two weeks. You can evaluate one workflow.

A good candidate workflow has these properties:

  • It runs at least a few times a week, so you get enough attempts
  • Somebody currently does it by hand and can tell you how long it takes
  • A wrong output is obvious rather than subtle
  • It does not touch money, contracts or anything you cannot undo

Weekly status summaries, first drafts of routine documents, sorting an inbound queue and checking a plan for internal contradictions all fit. Anything that ships to a customer without a human reading it does not.

Agree the number before you start

Write down, before the trial begins, the single number that decides it. Then write down the threshold. What it gets compared against comes from the numbers worth capturing before an agent goes live.

Useful shapes for that number:

  • Share of agent outputs the reviewer approved with no edits or light edits
  • Minutes of human time the workflow takes now versus during the trial
  • Number of items in a backlog cleared per week
  • Count of errors caught before the work left the team

Pick one. Two numbers turn into an argument about which one mattered, and picking the wrong one gives you a pilot that measures the wrong thing. And set the threshold as a specific figure your team chooses from its own baseline, not from a vendor's marketing. Measure the baseline in the week before the trial, by hand if necessary, otherwise you have nothing to compare against.

Write the stopping rule

Trials drift because nobody wrote down what failure looks like. It is easier to write once you have priced the likely failures in hours to correct rather than money. Fix that with a rule you can apply without a meeting.

Examples of a workable stopping rule: stop if the reviewer has to rewrite most outputs for three days running; stop if the workflow produces something that would have caused real harm had it not been caught; stop if nobody has used it for two days because it is easier to do the work by hand.

Stopping early is a good outcome. It saves the second week and gives you a clear answer.

The two weeks themselves

Keep the shape simple.

  • Days one and two: set up, connect the workflow, and have the reviewer run through a handful of outputs to calibrate what good looks like
  • Days three to eight: run it for real, with the reviewer logging every output as approved, edited or rejected in one line each
  • Days nine to twelve: no changes to the setup, so you get a clean stretch of data
  • Days thirteen and fourteen: the named decision maker reads the log, checks the number against the threshold, and writes a paragraph saying yes or no and why

The log is the whole trial. One line per output, three categories, plus a short note when something was rejected. Anything more elaborate will be abandoned by day four.

What to ignore during the trial

Some things feel like signal and are not. Enthusiasm in the first three days is not signal; new tools are fun. Volume of outputs is not signal either, since an agent that produces more work for a reviewer is producing cost, not value. Neither is the opinion of anyone who did not use it.

A note on approval

If you are trialling something like SicherOne, where a human approves agent output before it ships, the approval step is part of what you are measuring. Time spent reviewing counts as time spent. A tool that halves the drafting and doubles the checking has not helped anyone.

At the end of the fortnight, the named person writes one paragraph. Yes, and here is the number. Or no, and here is the number. That paragraph is worth more than any demo.

← All posts

We're building the future of community events and financial wellness

See how Eventify and WealthWise change the way people find events and manage money.

Get Started