AI, explained
How do I run an AI pilot?
Most AI pilots succeed and then die, because nobody agreed in advance what number would justify going to production.
Choose a process, not a technology
The candidate should have three properties: a cost you can count today (hours, queue length, error rate), data you already hold and can access this month, and a named person who owns the outcome. Pilots chosen because a technology is interesting produce demonstrations. Pilots chosen because a queue is too long produce decisions.
Measure the baseline before you build
This is the step that gets skipped and the one that decides everything afterward. Two weeks of the current process instrumented: volume, handling time, error rate, cost per item. Without it, every result is arguable, and the argument is always won by whoever is most senior rather than by the data.
Agree the threshold and the kill criterion
Write down, before building, what result justifies production and what result stops the work. For example: extraction accuracy above an agreed level on a held-out set, handling time down by a stated share, and cost per item below the current one. Get the person who controls the production budget to agree to those numbers in advance. A pilot without a kill criterion never ends, it just gets quieter.
Build against reality
- Real data, including the messy cases. A pilot on curated samples measures nothing.
- Real users, a small group, doing their actual work rather than testing.
- The production constraints from day one: access control, audit trail, and the human review queue. These are what turn a prototype into a system, and discovering them at the end is what kills pilots.
- A labelled evaluation set built during the pilot. It is the asset you keep whatever the decision.
Six to eight weeks, then decide
Long enough to see the awkward cases, short enough that the organisation has not moved on. At the end there are three honest outcomes: go to production, stop with a written reason, or extend once with a specific question to answer. Anything else is drift. Digiton runs pilots on this shape and starts them with a short AI audit that produces the baseline and the thresholds before a line of code is written.
Frequently asked questions
How do I run an AI pilot?
Choose one process with a countable cost today, measure the baseline for two weeks before building, agree a success threshold and a kill criterion in writing with the production budget holder, then build for six to eight weeks against real data, real users and production constraints. Decide on the number.
How long should an AI pilot run?
Six to eight weeks of build after a two week baseline measurement. Shorter and the awkward cases never appear. Longer and the sponsor changes role, the priority moves, and the pilot ends by evaporation rather than by decision. Extend once at most, with a specific question to answer.
Why do most AI pilots never reach production?
Three reasons. No baseline was measured, so improvement cannot be proven. No threshold was agreed, so success is a matter of opinion. And production constraints such as access control, audit trails and the human review queue were left until the end, where they look like a second project nobody funded.
Related
Ready to put AI to work?
Book a discovery audit and we will map the highest-ROI AI agents and automations for your business.
Book a discovery audit →