Predictive maintenance: from vibration sensor to a work order technicians will actually do
However good the anomaly detection model, it is useless if every alert becomes a maintenance order that nobody trusts.
In brief
- Too many false alarms is one of the leading reasons PdM programmes fail, so the FDE's job is to filter alerts before they reach the CMMS.
- Create a work order only once the risk is confirmed: the alert persists, a manual reading verifies it, and the machine is in zone C under ISO 20816.
- An alert must say why it fired, and repair outcomes must flow back into the system so the next alert is more accurate.
Picture the third week after a vibration monitoring system goes live. The plant’s maintenance planner has set up an email filter that sends every alert to a separate folder, and he no longer opens it. The dashboard is still full of green and red, the model still runs, but nobody acts on it any more.
That scenario matches two causes of failure that maintenance professionals often cite, and neither lies in the model. Reliable Magazine identifies excessive false alarms as one of the leading reasons PdM programmes fail. Programmes also tend to collapse when technicians and engineers are not brought into the rollout.
If you are an FDE sent to a plant, the work that decides success or failure is not in the notebook. It is in the step that turns a vibration reading into a work order a technician believes is worth doing. That step can be built, provided you take it one stage at a time and in the right order.
Collecting data is not a maintenance strategy
GroundUp points to a root error: many programmes mistake data collection for a maintenance strategy. Sensors go in, data flows, and the readings land on an engineer who is already overloaded and must interpret them alone, usually at the worst possible moment.
Your real product, then, is a chain of decisions, not a model. Every reading must pass three gates before it becomes work for someone else: is there an anomaly, has that anomaly been confirmed, and is the risk large enough to schedule a repair?
IVC Technologies puts the principle neatly: work orders are created from confirmed risk, not from preliminary signals.
Which threshold is right for which machine?
The first gate needs a threshold, and this is where many teams go wrong from the start. ISO 20816 is the current standard for evaluating mechanical vibration measured on the non-rotating parts of machines, with separate parts for different classes of machine.
James Otremba of Acoem USA stresses that ISO limits should be used as evaluation guidance, not as a pass/fail number applied to every machine.
A better approach combines two layers. The first is a threshold by machine type, because a pump and a large fan do not vibrate the same way. The second is a statistical or baseline alert, which sets limits around the actual behaviour of that particular machine.
ISO 20816 also gives you a very useful marker to map onto the maintenance process.
The standard divides vibration levels into zones A, B, C and D. Zone C means the machine is no longer fit for continuous operation but can run for a limited period until remedial action can be scheduled.
That description matches a planned work order almost perfectly.
Example: a cooling water pump
Try it with numbers. Suppose the wireless sensor on pump P-101 takes a reading every 10 minutes, or 144 times a day. If just 2% of readings cross the threshold because of noise, this one machine generates almost 3 alerts a day; across 20 machines that is nearly 58 alerts a day.
If every alert becomes a work order, the backlog will balloon within a week and the planner will set up the email filter from the opening paragraph. IVC Technologies recommends clear rules for when an alert becomes a work request, precisely to avoid a swollen backlog and an overloaded planner. Here is a simple version of such a rule:
def decide(asset, readings, route_check):
limit = asset.baseline_limit # learned from this machine's own history
recent = readings[-6:] # last 6 readings, about 1 hour
over = [r for r in recent if r.velocity > limit]
if len(over) < 4:
return "WATCH" # transient: log it, create no work
if route_check is None:
return "REQUEST_ROUTE" # ask a technician to take a manual reading next shift
if route_check.confirms and route_check.zone == "D":
return "ESCALATE" # never fall back to WATCH: alert the responsible person now
if route_check.confirms and route_check.zone == "C":
return "PLANNED_WO" # confirmed risk: schedule the repair
return "WATCH"
This code is designed carefully only for zone C, because that is the zone that maps naturally onto a planned work order.
The zone D branch does just one thing: it stops a confirmed reading in the most severe zone from quietly reverting to WATCH, and hands it straight to the responsible engineer to decide.
The REQUEST_ROUTE step is not redundant. IVC Technologies argues that hybrid programmes, combining wireless sensors with manual route-based readings, help avoid creating work orders for transient or unimportant conditions. The technician who takes the reading also has reason to trust the result, because they confirmed it themselves.
Alerts must explain themselves
Even after passing all three gates, an alert has one more step before it reaches a technician: it must be written so the reader understands it at once.
Reliable Magazine observes that maintenance teams rarely trust black-box models, and LLumin argues that an alert without context is just noise: to be actionable, anomaly detection must be explainable.
A work order for P-101 should read like this:
P-101, motor-side bearing. Vibration exceeded the machine’s own baseline in 4 of 6 readings over the past hour, with a rising trend. Manual reading on the morning shift confirmed zone C under ISO 20816. Recommend inspecting the bearing at the next planned shutdown.
When closing the job, the technician picks one of three codes: fault found as predicted, different fault found, or nothing found.
LLumin describes feeding this outcome data back into the alerting system as the way to make later alerts more accurate. The “nothing found” code is the most valuable, because it tells you which machine’s baseline is set too tight.
Step by step at the customer site
The first task is to sit down with the technicians and the maintenance planner before writing a line of code. Ask them which machines fail often, how they currently take manual route readings, and how many work requests they can handle in a week. That last number is your alert budget.
Next, group assets by machine type so the appropriate part of ISO 20816 can be applied as guidance. Let the system learn each machine’s baseline over a period of stable operation. Then write the rules for turning alerts into work requests together with the planner, so that they are a co-owner rather than the one left holding the consequences.
The ESCALATE branch also needs to be agreed in writing with the customer, because the code only forwards the alert; it does not decide on anyone’s behalf. Agree on three things: who receives zone D alerts on each shift, how quickly they must respond, and who has the authority to stop the machine. Leave those three blank and the most serious alerts are the ones most likely to fall into a void.
Finally, lock down the alert template and the set of outcome codes in the CMMS before go-live. If the customer has no field for recording outcomes, propose adding one from the outset, because without it the feedback loop never closes.
The mistakes that discredit the whole programme
The most common mistake is switching on alerts for every machine at once with the same ISO threshold. Next comes pushing alerts straight into the CMMS without a confirmation step. Then sending a bare number with a chart and waiting for engineers to work it out.
All three mistakes cost more than they appear to. GroundUp warns that trust, once lost, is very hard to regain. So start with a few critical machines, with few alerts but every one of them right, and only then expand.
This is also a skill worth making visible when you look for a job. If you are aiming for an FDE role in industry, read job descriptions carefully for mentions of CMMS, IIoT or vibration analysis, then describe in your CV a time you turned signals into work that operators actually followed.
The measure of a PdM deployment is not accuracy on a test set. It is whether, six months later, the planner still opens the alert emails.
Was this article useful?
Thanks for the feedback!
6 sources
- Getting Predictive Maintenance Right with Grounded AI and IIoT Practices (Reliable Magazine)
- Integrating Vibration Monitoring into CMMS Systems (IVC Technologies) · 2026-01-17
- Setting Better Vibration Alarm Limits By Machine Type (Acoem USA, James Otremba) · 2026-08-27
- ISO 20816 Vibration Severity Zones: A, B, C and D Explained (Fabrico) · 2026-07-07
- How to Build Technician Trust in AI-Powered Alerts (LLumin) · 2026-03-11
- Why Sensor Maintenance Programs Fail in Factories (GroundUp) · 2026-01-13