Predictive maintenance planner
A planning workflow that combines condition signals, failure history, work orders, parts and operating windows to rank equipment risk for human approval.
The planner ranks maintenance risk without making safety decisions
A predictive maintenance planner combines condition signals, failure history, work orders, parts availability and operating windows. From those it ranks equipment risk and proposes work, and a person reviews every proposal.
It does not declare equipment safe and it does not override controls. Authorised staff keep responsibility for diagnosis, work scope, isolation and return to service. Good candidates have costly downtime, a planning backlog, relevant condition data and consistent records in the computerised maintenance management system (CMMS). Keep fixed intervals for regulated tasks, and for failure modes that give no useful warning signal.
Planning capacity is easier to observe than avoided downtime
A ranked queue reduces the time planners spend gathering facts, checking parts and drafting work orders. That capacity is easy to observe in a time study, but it is not a financial benefit by itself. Count money only when the capacity avoids hiring, overtime or vendor spend, or supports extra output at a proven constraint.
Avoided downtime is the larger prize and the harder attribution. A claim needs a chain: the recommendation, the failure pattern it pointed at, the approved action, and comparable events that show what the failure would have cost, with allowance for other plausible causes. An alert on its own has no value.
Measure operating behaviour with quality guardrails
| KPI | What it shows | Measurement approach |
|---|---|---|
| Engineering-attributed avoided downtime estimate | Probable interruption avoided | Engineering review of comparable failures |
| Planner hours per weekly schedule | Capacity released | Pilot time study |
| Emergency work-order share | Reactive work trend | Comparable CMMS cohort |
| Recommendation acceptance | Advice used in approved work | Accepted, modified and rejected counts |
| False-positive rate | Unnecessary work | Alerts with no actionable finding |
| Recall and missed-failure rate | Relevant failures detected or missed | By failure family and asset criticality |
| Alert latency | Time available to act | Signal timestamp to planner availability |
| Schedule adherence | Approved plans completed | Work completed in its planned window |
Segment false positives by asset and failure mode, and give missed failures the same attention, because a model that misses quietly costs more than one that alerts too often.
The value model keeps capacity outside the P&L
Estimate the capacity pool first, as a realisation-adjusted proxy:
Gross planner capacity =
planners × hours saved weekly × 52 × loaded hourly rate
Realisation-adjusted operational capacity proxy =
gross planner capacity × capacity realisation factor
The proxy describes the scale of the capacity, not P&L value. A planner hour becomes money through what it enables: removed spend at the planner's cost, or downtime avoided at the plant's contribution margin, which is far above any hourly rate. Only evidenced avoided hiring, overtime or vendor spend counts on the cost route. Downtime is calculated separately:
Attributed downtime value =
engineering-attributed avoided downtime hours
× contribution margin per constrained hour
× downtime attribution factor
Use contribution margin per constrained hour, allow for other causes through the attribution factor, and never count the same loss twice.
An illustrative multi-plant scenario
These assumptions are fictional. They are not a benchmark, a forecast, a guarantee or an Epicube quote.
| Input | Fictional assumption | Evidence needed |
|---|---|---|
| Planners and reliability engineers | 24 | Named users |
| Planning time reduction | 14 hours each week | Pilot time study |
| Loaded planner rate | 480 SEK per hour | Finance-approved cost |
| Avoided downtime estimate | 18 hours per year | Comparable failures |
| Contribution per constrained hour | 100,000 SEK | Finance and production records |
| Capacity realisation | 40% | Approved capacity use |
| Downtime attribution | 40% | Engineering and finance review |
| First-year cost | 1.35 MSEK build plus 0.35 MSEK operation | Scoped supplier and operating costs |
Gross planner capacity = 24 × 14 × 52 × 480 = 8.39 MSEK
Realisation-adjusted operational capacity proxy = 8.39 × 40% = 3.35 MSEK
This 3.35 MSEK is an operating pool, not P&L value.
Gross downtime value = 18 × 100,000 = 1.80 MSEK
Attributed downtime value = 1.80 × 40% = 0.72 MSEK
Counted annual financial benefit = 0.72 MSEK
Net first-year value = 0.72 - 1.70 = -0.98 MSEK
Illustrative payback = 1.70 ÷ (0.72 / 12) = about 28 months
The financial case counts only the attributed downtime. Capacity value is added later, and only if named costs are actually avoided.
A risk score becomes draft work only after operational checks
Connect the CMMS, the sensor platform, the asset registry, parts and production windows through consistent codes. Model inputs can include trends, threshold durations, operating regime and recent maintenance, and engineering rules apply the exclusions, statutory intervals and criticality thresholds on top.
Before a recommendation becomes draft work, the planner checks parts, skills, permits, windows and data freshness, and suppresses recommendations built on stale data. Define what happens during a sensor or CMMS outage, the maximum acceptable alert latency, and the fallback to existing rules. An outage must never read as healthy equipment.
Each recommendation shows its evidence and its uncertainty. An accepted recommendation becomes a draft work order in the CMMS, never a released instruction, and the record keeps the model version, the inputs and the decision.
Evaluate by time, using only the information that was available before each prediction. Keep each failure episode inside one data split and reserve a final untouched test period. Measure recall and the missed-failure rate by failure family and criticality, along with precision, warning lead time, calibration, false positives and latency. Replay the parts and operating windows that were available at the time, so the evaluation does not assume options the planner never had. Engineering signs off before the system drafts anything.
In operation, monitor missing sensors, timestamp errors, calibration changes, asset remapping, alert volume, acceptance and outcomes. Retraining needs reviewed evidence and a versioned evaluation.
Start with one failure family and clear sign-off
Start with one plant, one asset class and one costly failure family. Clean the identifiers, define the failure labels, baseline planner effort and downtime, then run in shadow mode.
Engineering and safety owners approve the assets, exclusions, thresholds and escalation rules. Existing lockout, permit, inspection and management-of-change processes stay exactly as they are, and no recommendation may suppress an alarm or delay mandatory maintenance.
Do not build when the failure history is unreliable, when sensors do not see the degradation, when asset identifiers conflict, when preventive work is already being deferred, or when planners have no room to act on recommendations. A threshold rule or a vendor tool is often enough. Stop if the likely avoided loss cannot justify the integration and governance.
Sources and methodology
The planning workflow, the attribution rules and the ROI model are Epicube analysis, and the worked case is illustrative. These independent Swedish research sources inform the industrial context:
- RISE: smart prediktivt underhåll covers Swedish research and decision support for predictive maintenance
- Vinnova: Future AI-based Maintenance documents a Swedish research programme on prediction, uncertainty and maintenance planning
