Accuracy is not the constraint

Most predictive maintenance programmes that fail do not fail on accuracy. They produce reasonable predictions that nobody acts on, and are quietly retired.

Being told an asset will fail is a request to do something expensive and disruptive on someone else’s authority. Whether that request is granted depends almost entirely on properties of the system that are not the model.

1. What the engineer is actually being asked

To take a working asset out of service, on a prediction, against a schedule already agreed with operations, with the cost in their budget and the consequence of being wrong on them.

Against that, “the model says so” is a weak argument — correctly so. The engineer has context the model does not: this unit reads high after a restart, this sensor has drifted since March, this pattern appeared last winter and meant nothing.

The design goal is not to overrule that knowledge. It is to give it something to work with.

2. What makes a prediction actionable

Contributing factors, ranked. Not a score — which readings drove it, and by how much. That is evaluable against what the engineer knows.

A comparison to normal. Normal for this asset, in this duty, at this time of year. A fleet-wide average flags every unit in the coldest month and teaches people to ignore it.

A time horizon. “Elevated risk” is unusable. “Likely within two to six weeks” can be planned into a window that already exists.

A recommended action with a cost. Inspect, monitor, or intervene. Most alerts should resolve to “look at this on the next planned visit”. A system that only ever escalates gets muted.

A recorded route to disagree. Engineer overrides are the most valuable feedback available. If disagreement happens in a phone call, it is lost.

3. False positives are the whole game

Alert fatigue is not a nuisance, it is the failure mode. A system that cries wolf is abandoned within weeks, and the abandonment is permanent — a second attempt inherits the reputation of the first.

A conservative model that flags less and is right more is worth more than a sensitive one, at least initially. Trust is built by being right about small things before anyone accepts a large call.

Set the threshold with the maintenance team, in terms of how many alerts a week they can genuinely investigate, and treat that as a hard constraint rather than a tuning parameter.

4. The foundations decide the ceiling

Most of this work is not modelling. It is discovering that:

  • The same asset appears three ways across four systems.
  • Maintenance history is free text in work orders.
  • The sensor tag list was last accurate in 2019.
  • Failures were closed as “repaired” with no recorded cause.
  • Condition data and work history have no reliable join.

Without a trustworthy record of what failed and why, there is nothing to learn from. This is the data strategy problem in another sector, and it is where the time goes.

Failure coding is the highest-return fix. A short, enforced cause list at closure — chosen by the person who did the work, at the moment they did it — is worth more to a future model than any amount of additional sensing.

5. Start deterministic

The rules-based version is not a stepping stone to be embarrassed about. It is frequently the product:

  • Runtime hours since last intervention.
  • Readings outside a band agreed with engineering.
  • Rate of change, which catches more than absolute thresholds.
  • Inspection findings escalating over successive visits.

It is transparent, arguable, and it produces the instrumentation and the trust that any later model will need. Introduce a model only where it demonstrably beats this on the same instrumentation — and be prepared for it not to, on many assets.

6. Measure the thing you are trying to change

Availability is the headline number and, in most organisations, one nobody can reproduce, because the definition moves depending on who is counting.

The same outage has three durations: fault-to-repair, alarm-to-return-to-service, and production loss. Maintenance reports one, operations another. Both are right; they are answering different questions with the same word.

Settle it explicitly:

  • One written definition per metric — what starts the clock, what stops it, what is excluded.
  • Planned and unplanned always separate. Combining them hides the only distinction that matters for improvement.
  • The clock started by an event, not a person. Otherwise you are measuring reporting behaviour.
  • Exclusions fixed in advance — weather, grid constraint, third party — rather than extended when a month looks bad.
  • Restatement allowed and visible. Numbers improve as investigation completes; a figure that quietly changed is worse than one that changed with a note.

The usual surprise once this holds still: more downtime is spent waiting — for a part, a permit, an engineer, a window — than on repair. Nobody sees it while the metric is repair duration, and it is often the cheapest thing to fix.

7. Competency is part of the same picture

A prediction is worthless if the person who could act on it cannot legally do the work. Certification expiry is an operations problem, not an HR one: it belongs in the same view as the plan.

That means certifications held against the person with type, expiry and evidence; work types that declare what they require; availability that accounts for it, so someone whose card expires on the 14th is not allocatable on the 15th; and alerts to the planner weeks out rather than to a compliance mailbox.

8. Monitoring the model itself

  • Watch inputs as well as outputs. A distribution shift usually arrives first and is an alert in its own right.
  • Watch the override rate. A rise means the people closest to the work have stopped trusting it, and they are usually right.
  • Define and test the rollback. An untested rollback is a plan, not a control.

A sequence that works

  1. Resolve asset identity. Nothing works before this.
  2. Enforce failure coding at closure, and wait. You are building the training set.
  3. Settle the availability definitions, so improvement is measurable.
  4. Ship the deterministic version, fully instrumented.
  5. Introduce a model only where it beats the rules, with the same explanation surface.
  6. Review overrides quarterly, not just accuracy.

We build systems that show their working — RailGard AI scores with every contributing factor logged, and ECS Group Command holds the asset register, inspection status and certification that a model of this kind needs underneath it.


Getting operational forecasting into production? Get in touch, or read about our energy and utilities work.

← Back to Energy & Utilities