Predictive maintenance planner
A maintenance planning workflow that combines condition signals, failure history, work orders, parts and operating windows to rank equipment risk for human approval.
The planner ranks maintenance risk without making safety decisions
A predictive maintenance planner combines condition signals, failure history, parts and operating windows. It ranks asset risk and proposes work for human review.
It does not declare equipment safe or override controls. Authorised staff retain responsibility for diagnosis, work scope, isolation and return to service. Good candidates have costly downtime, planning backlogs, relevant condition data and consistent computerised maintenance management system (CMMS) records. Keep fixed intervals for regulated tasks and failures without useful warning signals.
Planning capacity is easier to observe than avoided downtime
A ranked queue can reduce time spent gathering facts, checking parts and drafting work. This capacity is easier to observe through time studies, but it is not a realised financial benefit. Count financial value only when it avoids hiring, overtime or vendor spend, or supports extra output at a proven constraint.
Avoided downtime is harder to attribute. Link the recommendation to a failure pattern, approved action and comparable events, allowing for other plausible causes. Alerts alone have no value.
Measure operating behaviour with quality guardrails
| KPI | What it shows | Measurement approach |
|---|---|---|
| Engineering-attributed avoided downtime estimate | Probable interruption avoided | Engineering review of comparable failures |
| Planner hours per weekly schedule | Capacity released | Pilot time study |
| Emergency work-order share | Reactive work trend | Comparable CMMS cohort |
| Recommendation acceptance | Advice used in approved work | Accepted, modified and rejected counts |
| False-positive rate | Unnecessary work | Alerts with no actionable finding |
| Recall and missed-failure rate | Relevant failures detected or missed | By failure family and asset criticality |
| Alert latency | Time available to act | Signal timestamp to planner availability |
| Schedule adherence | Approved plans completed | Work completed in its planned window |
Segment false positives by asset and failure mode. Track missed failures equally closely.
The value model keeps capacity outside P&L
Estimate the operating pool as a realisation-adjusted operational capacity proxy:
Gross planner capacity =
planners × hours saved weekly × 52 × loaded hourly rate
Realisation-adjusted operational capacity proxy =
gross planner capacity × capacity realisation factor
The proxy describes capacity scale, not P&L value. Count only evidenced avoided hiring, overtime or vendor spend. Calculate downtime separately:
Attributed downtime value =
engineering-attributed avoided downtime hours
× contribution margin per constrained hour
× downtime attribution factor
Use contribution margin, allow for other causes and avoid counting the same loss twice.
An illustrative multi-plant scenario
These fictional assumptions are not a benchmark, forecast, guarantee or Epicube quote.
| Input | Fictional assumption | Evidence needed |
|---|---|---|
| Planners and reliability engineers | 24 | Named users |
| Planning time reduction | 14 hours each week | Pilot time study |
| Loaded planner rate | 480 SEK per hour | Finance-approved cost |
| Avoided downtime estimate | 18 hours per year | Comparable failures |
| Contribution per constrained hour | 100,000 SEK | Finance and production records |
| Capacity realisation | 40% | Approved capacity use |
| Downtime attribution | 40% | Engineering and finance review |
| First-year cost | 1.35 MSEK build plus 0.35 MSEK operation | Scoped supplier and operating costs |
Gross planner capacity = 24 × 14 × 52 × 480 = 8.39 MSEK
Realisation-adjusted operational capacity proxy = 8.39 × 40% = 3.35 MSEK
This 3.35 MSEK is an operating pool, not P&L value.
Gross downtime value = 18 × 100,000 = 1.80 MSEK
Attributed downtime value = 1.80 × 40% = 0.72 MSEK
Counted annual financial benefit = 0.72 MSEK
Net first-year value = 0.72 - 1.70 = -0.98 MSEK
Illustrative payback = 1.70 ÷ (0.72 / 12) = about 28 months
The financial case counts only attributed downtime. Add capacity value later only if named costs are avoided.
Risk scores become draft work only after operational checks
Connect the CMMS, sensor platform, asset registry, parts and production windows through consistent codes. Inputs can include trend, threshold duration, operating regime and recent maintenance. Engineering rules apply exclusions, statutory intervals and criticality thresholds.
Before drafting, check parts, skills, permits, windows and data freshness. Suppress stale-data recommendations. Define sensor and CMMS outage behaviour, maximum alert latency and fallback to existing rules. An outage must never imply healthy equipment.
Show evidence and uncertainty. Accepted recommendations become draft CMMS work orders, never released instructions. Record the model version, inputs and decision.
Evaluate by time, using only information available before each prediction. Keep each failure episode in one split and reserve a final untouched test period. Measure recall and missed-failure rate by failure family and criticality, precision, warning lead time, calibration, false positives and latency. Replay the parts and windows available then. Engineering must sign off before drafting.
Monitor missing sensors, timestamp errors, calibration changes, asset remapping, alert volume, acceptance and outcomes. Retraining requires reviewed evidence and versioned evaluation.
Start with one failure family and clear sign-off
Start with one plant, asset class and costly failure family. Clean IDs, define labels, baseline planner effort and downtime, then run in shadow mode.
Engineering and safety owners approve assets, exclusions, thresholds and escalation rules. Existing lockout, permit, inspection and management-of-change processes remain. Recommendations must not suppress alarms or delay mandatory maintenance.
Do not build when history is unreliable, sensors miss the degradation, IDs conflict, preventive work is deferred or planners cannot act. Prefer an adequate threshold rule or vendor tool. Stop if likely avoided loss cannot justify integration and governance.
Sources and methodology
The planning workflow, attribution rules and ROI model are Epicube analysis. The worked case is illustrative. These independent Swedish research sources inform the industrial context:
- RISE: smart prediktivt underhåll covers Swedish research and decision support for predictive maintenance
- Vinnova: Future AI-based Maintenance documents a Swedish research programme on prediction, uncertainty and maintenance planning
