A pilot that works is a strange thing to celebrate, because the pilot was never the hard part. Getting to a demo requires a model, a dataset and someone motivated. Getting to production requires everything else.

Most organisations we speak to have at least one pilot that never crossed that gap. The reasons repeat.

The pilot was measured on the wrong thing

Pilots are judged on model performance. Production is judged on whether the number that needed to move, moved.

Those come apart more often than people expect. A model can be materially better than the current process and still change nothing, because the decision it informs was not the constraint. If nobody asked which number was supposed to move before the work started, there is no way to answer whether it did.

It was built beside the workflow, not inside it

The pilot lives in a notebook, a dashboard or a separate tool. Using it requires somebody to stop what they are doing, go somewhere else, and bring the answer back.

That works for the enthusiast who built it. It does not survive a busy week.

The best outcome for most operational AI is that nobody notices it: a field that fills itself in, a case that arrives pre-prioritised, a pack that assembles itself. A separate tool nobody opens delivers nothing, however good the model is — and adoption problems are usually design problems wearing a different hat.

Nobody owned it after the pilot

A pilot has a champion. Production needs an owner, which is a different and less appealing job: monitoring, retraining, handling the cases the model gets wrong, answering questions about it a year later.

If that ownership is not agreed before launch, the system decays quietly. Data drift is the usual mechanism — the world changes, the model does not, and performance degrades slowly enough that nobody notices until it is obviously wrong.

The governance was left until last

In regulated environments this is where most stalls actually happen, and it is avoidable. If explainability, audit trails, access control and human oversight are treated as a compliance step after the build, the answer is often that it cannot go live as built.

Bringing that conversation forward costs very little at design time and saves the project. It is the same argument as accessibility: retrofitting is always more expensive and usually worse.

What crossing the gap actually takes

  • Name the number the work is supposed to move, before anything is built.
  • Put it where the work already happens, not beside it.
  • Ship it in production shape from the first sprint — tested, observable, documented, deployable.
  • Decide who owns it in six months, and what they are expected to do.
  • Design the governance in, not after.
  • Be willing to conclude it is not worth building. Some of the most valuable work is finding that out cheaply.

That last one is uncomfortable to sell and it is often the right answer. We would rather talk a client out of a model they do not need than deliver one nobody uses.

The question to ask about your own pilot

Not “did it work?” but:

If this ran every day for a year with nobody watching it, what would go wrong, and who would notice?

If there is no answer, the gap to production is the answer.


Sitting on a pilot that has not moved? Get in touch, or see how we put AI into production.

← Back to the blog