107 Predictive Maintenance: Signals and Vibration
107.1 Overview
This first route connects maintenance strategy and business evidence to a signal pipeline, then develops vibration features into reviewable fault evidence.
This is part 1 of 2. Continue with Predictive Maintenance: Models and Rollout for the second focused route.
107.2 Start With the Story
Turn a Weak Signal Into a Planned Check
Picture a motor that still runs but sounds different each morning. Heat, sound, current, and vibration may give early clues. Predictive maintenance uses such clues to plan a check before a failure, without replacing parts just because a date arrived.
Begin with one failure you can name. Choose a signal that can change before that failure. Record the machine state, load, speed, place, and time beside each reading. A change during heavy work may be normal; the same change at light work may matter.
Set an action before setting an alert. Decide who reviews it, how soon they respond, and what evidence confirms the fault. Start with a small group of machines. Count useful warnings, false alarms, missed faults, saved stops, and needless work.
Follow one machine through a full review. Record its normal sound. Record its normal heat. Record its normal current. Keep the work state. Keep the speed. Keep the load. Mark each repair. Mark each part change. Mark each sensor change. Compare like work with like work.
When a clue moves, ask simple questions. Is the reading fresh? Is the sensor sound? Did the work state change? Did the room change? Did the speed change? Did a new part change the base line? Can a second check see the same shift? Is the harm large enough to stop now?
Write the response before the next alert. Name the reviewer. Name the time limit. Name the safe first check. Name the proof needed for repair. Name the proof needed to keep running. Make a missed response visible. Do not let a warning sit in a chart with no owner.
Review the result after the work. Was there a real fault? Was the warning early enough? Was any good part replaced? Was a fault missed? Did the check add risk? Did the repair change the signal? Use those answers to keep, change, or remove the rule.
Scale only after the small group works. Add one machine type at a time. Keep a return path. Train the people who receive the warning. Check every shift. Check quiet and busy work. Check heat and cold. Check after power loss. Keep the old method until the new path earns trust.
The simple story has a limit. A pattern can show that something changed; it may not prove why. Safe action still needs physical checks and maintenance judgment. It also needs a clear stop rule when the evidence is weak.
Use Practitioner to build the signal-to-work path. Use Under the Hood to study vibration, models, cost, and error limits.
Start with a machine that has not failed yet, but is beginning to leave evidence in vibration, temperature, current, sound, or operating history. Predictive maintenance is the story of turning weak signals into a justified intervention before downtime, safety risk, or unnecessary replacement costs appear.
107.3 Learning Objectives
After completing this chapter, you will be able to:
- Apply predictive maintenance patterns using IoT sensor data
- Compare reactive, preventive, and predictive maintenance strategies
- Design vibration analysis systems for rotating machinery
- Implement machine learning models for remaining useful life prediction
- Calculate ROI for predictive maintenance investments
This chapter moves from maintenance problem to deployable IIoT program:
- First we compare reactive, preventive, and predictive maintenance so the value case is clear.
- Then we design the signal chain around named assets, fault signatures, operating context, and reviewable evidence.
- Next we inspect vibration, thermal, and machine-learning methods before turning alerts into work orders.
- Finally we test the economics, implementation phases, quizzes, and common pitfalls that decide whether technicians trust the system.
Checkpoints recap the operating decisions; deep calculations and interactives let you inspect the numbers without losing the main path.
Predictive maintenance is like visiting the doctor for a check-up before you feel sick. Instead of waiting for a factory machine to break down (which is expensive and dangerous), sensors listen to the machine’s vibrations, temperature, and sounds to spot tiny warning signs weeks in advance. It is the same idea as your car telling you to change the oil at 5,000 miles instead of waiting for the engine to seize — except IoT sensors do it automatically, 24 hours a day.
107.4 Prerequisites
Before diving into this chapter, you should be familiar with:
- Industry 4.0 Fundamentals: Core concepts of Industry 4.0 and smart manufacturing
- Real-Time Requirements and ISA-95: Understanding of automation levels and data flows
- Data Storage and Databases: Time-series storage for industrial data
Learning Resources:
- Content Hub - Search for predictive maintenance practice checks, ROI calculations, vibration analysis, and ML model selection resources
- Simulations Hub - Explore vibration signal analysis simulations and FFT visualization tools
- Videos Hub - Watch real-world predictive maintenance and smart-factory implementation examples
- Knowledge Gaps Hub - Address common misconceptions about maintenance strategies and ROI calculations
- Concept Navigator - Explore how predictive maintenance connects to IIoT, ML/AI, time-series databases, and digital twins
107.5 PdM Turns Signals into Work
Predictive maintenance is valuable when a signal changes a maintenance action before the asset fails. The goal is not to collect every possible waveform. The goal is to detect a developing fault early enough to order parts, schedule labor, protect safety, and avoid unplanned downtime.
The strongest candidates are assets with expensive failure modes and measurable degradation patterns: motors, pumps, compressors, fans, gearboxes, spindles, conveyors, chillers, turbines, and critical bearings. The signal may come from vibration, temperature, acoustic emission, motor current, pressure, oil debris, lubricant analysis, cycle counts, or control-system operating context.
Ground pdm turns signals into work with the visual at Figure 107.1. Start from Predictive Maintenance Pipeline, but keep Vibration visible while evaluating condition-to-action path: predictive maintenance has to preserve the link from physical condition signal to maintenance action, not just display a.
Compare Predictive Maintenance Pipeline with Vibration inside the visual at Figure 107.1. Next find Sensor data, which completes the scope of condition-to-action path: predictive maintenance has to preserve the link from physical condition signal to maintenance action, not just display a. The decision in pdm turns signals into work must preserve that labelled boundary. A useful deployment starts with one named asset and one named fault. For example, a drive-end bearing on a cooling pump might use a triaxial accelerometer, motor current, and discharge pressure. The first proof is whether the system can distinguish normal load changes from a developing bearing defect, then raise an alert early enough for a planned inspection. If the alert cannot become a work order with parts, labor, and safe access, the analytics result has not yet created maintenance value.
The business case should also include what will not be predicted. Some faults arrive too suddenly, some assets lack repeatable operating cycles, and some failures are cheaper to repair after failure than to instrument continuously. A strong PdM program is selective: monitor the expensive, repeatable, observable failure modes first, then expand when technicians trust the alerts and the measured downtime reduction is real.
- Asset criticality: Start where failure stops production, damages equipment, creates safety risk, or causes high emergency repair cost.
- Fault signature: Choose sensors because they expose a known degradation mechanism, not because they are easy to install.
- Maintenance action: Define the work order, part, inspection, slowdown, shutdown, or operator response that the alert should trigger.
107.6 Design the Signal Chain First
A practical PdM pilot starts with failure modes and effects analysis, asset history, and maintenance records. If bearing wear is the target, the sensor placement, sampling rate, mounting method, machine speed, load state, and baseline period matter more than the dashboard. A poorly mounted accelerometer can produce clean-looking data that is useless for diagnosis.
Useful features depend on the asset. Rotating equipment may use RMS vibration, peak acceleration, crest factor, velocity, FFT bands, envelope spectra, bearing defect frequencies such as BPFO and BPFI, temperature trend, and motor current signature analysis. Process equipment may use pressure differential, flow, valve position, energy use, startup time, or compressor discharge temperature.
For a first pilot, collect enough context to make each alert reviewable by a maintenance engineer. Record the sensor location, axis orientation, mounting method, sampling rate, machine speed, load, product recipe, recent repair history, alarm threshold, and who inspected the asset after the alert. Then compare the alert to a physical finding: looseness, imbalance, misalignment, cavitation, lubrication failure, clogged filter, damaged bearing race, or no fault found. This evidence record keeps the project out of demo mode and gives technicians a reason to trust or challenge the model.
- Pick the failure mode. Name the component, mechanism, and business consequence: bearing wear, imbalance, misalignment, cavitation, lubrication loss, clogged filter, or belt degradation.
- Capture operating context. Store speed, load, recipe, product, duty cycle, ambient temperature, maintenance event, and asset state with the sensor values.
- Choose the detection method. Start with thresholds and trend rules where physics is clear; add anomaly detection, random forest, XGBoost, or sequence models only when labeled history and drift controls justify them.
- Close the workflow. Route alerts to CMMS/EAM systems such as IBM Maximo, SAP PM, or maintenance ticketing, and track whether technicians confirmed the fault.
107.7 PdM Needs Physics and Workflow
Condition data needs a traceable path from sensor to action. An IEPE accelerometer, MEMS sensor, current transformer, oil particle counter, or temperature probe may connect to an edge DAQ, PLC, or gateway. The gateway may compute FFT features locally, forward time-series data through OPC UA or MQTT, and store results in a historian, AVEVA PI System, InfluxDB, TimescaleDB, or cloud data lake.
Sampling choices change what faults can be seen. Nyquist limits, anti-alias filtering, window length, spectral resolution, tachometer references, sensor orientation, mounting stiffness, and unit conversion all affect the signal. Downsampling a 10 kHz vibration waveform to a 1 Hz dashboard trend can erase the bearing information that maintenance needed.
Model operations are part of the design. Alerts need severity, confidence, asset id, feature values, baseline comparison, recent maintenance state, and recommended inspection. Teams must handle false positives, false negatives, seasonal operation, new product mixes, replaced parts, firmware changes, sensor drift, and concept drift. A PdM system that cannot learn from technician outcomes becomes an expensive alarm list.
The data model should keep raw windows, derived features, and maintenance events connected. A feature such as vibration RMS or envelope energy should carry the source sensor, asset id, timestamp, window size, filter settings, firmware version, and operating state. The work-order outcome should then link back to the same feature window. That linkage lets the team retrain models, audit false alarms, compare edge firmware versions, and prove whether the warning arrived before the spare-part lead time and planned maintenance window.
- Feature lineage: Keep sampling rate, window, filter, unit, sensor location, and firmware/configuration version with derived features.
- Alert economics: Tune thresholds against downtime cost, inspection cost, spare-part lead time, and technician trust.
- Feedback loop: Capture technician disposition, part condition, root cause, and return-to-service date so the model and maintenance plan improve.
Checkpoint: Signal Chain Design
You now know:
- Predictive maintenance starts with one named asset and one named fault, not with every possible waveform.
- Reviewable alerts need sensor location, axis orientation, sampling rate, speed, load, baseline period, and recent maintenance state.
- The data model should keep raw windows, derived features, and work-order outcomes connected so false alarms and confirmed faults can improve the system.
107.8 Introduction
One of the highest-value applications of Industrial IoT is predictive maintenance. By continuously monitoring equipment health through vibration, temperature, and other sensors, manufacturers can detect failures weeks before they occur, scheduling repairs during planned downtime rather than suffering costly unplanned outages.
Core Concept: Predictive maintenance uses condition signals such as vibration, temperature, acoustic, current, oil, pressure, and operating context to detect degradation early enough for planned maintenance. Why It Matters: The value comes from converting uncertain failure risk into scheduled work. A useful program reduces emergency repair, lost production, safety exposure, and wasted preventive replacement, but the economics depend on the asset, failure mode, spare-part lead time, labor availability, and measured alert quality. Key Takeaway: Start with high-criticality assets and known fault signatures. Define the sensor, sampling rate, baseline period, alert threshold, work-order path, and evidence record before claiming a prediction window or ROI.
Hey there, young engineer! Let’s learn about predictive maintenance with the Sensor Squad!
Temperature Terry has a new job at a candy factory! His mission: keep the big machines running so they can make chocolate bars all day long.
The Problem: The giant chocolate mixer broke down yesterday! Now there’s no chocolate, and everyone is sad. The repair took 3 days because nobody knew it was about to break.
Sammy’s Solution: Be a Machine Doctor!
Sammy decides to become like a doctor who listens to your heartbeat. But instead of a stethoscope, Sammy uses special sensors:
- Vibration Sensor (like feeling a cat purr): Sammy sticks to the mixer and feels how it shakes. If it starts shaking funny, something’s wrong!
- Temperature Sensor (like checking for a fever): If the mixer gets too hot, it might be getting sick
- Sound Sensor (like hearing a squeaky wheel): Machines make different sounds when they’re healthy vs unhealthy
How Sammy Saves the Day:
- Monday: Sammy notices the mixer is shaking a tiny bit more than usual
- Tuesday: The shaking gets worse, and the temperature goes up a little
- Wednesday: Sammy sends an alert: “Hey! Fix me this weekend before I break!”
- Saturday: The maintenance team replaces a worn bearing in just 2 hours
- Monday: The mixer is back to making chocolate perfectly!
The Magic: Instead of waiting for the machine to break (and losing 3 days of chocolate!), Sammy helped fix it during the weekend when nobody needed it anyway. That’s called predictive maintenance - predicting problems before they happen!
Sensor Squad Memory Trick:
- Vibration = Feeling the machine’s “heartbeat”
- Temperature = Checking for “fever”
- Prediction = Being a fortune teller for machines
- Maintenance = Giving machines their medicine before they get really sick
107.9 Maintenance Strategies Comparison
Key Concepts
- Asset Criticality: Ranking equipment by production impact, safety consequence, repair cost, spare-part lead time, and whether failure stops a constrained process.
- Fault Signature: A measurable condition pattern, such as 1x vibration for imbalance, 2x vibration for misalignment, BPFO/BPFI bearing bands, rising temperature, current imbalance, or pressure drift.
- Baseline Profile: A reference record of normal behavior under known speed, load, recipe, environment, and maintenance state.
- Feature Extraction: Converting raw windows into reviewable values such as RMS vibration, crest factor, FFT bands, envelope energy, thermal rise, or motor-current harmonics.
- Remaining Useful Life (RUL): An estimate of time or cycles until a failure threshold is likely, valid only for the modeled fault mode and operating context.
- CMMS/EAM Integration: Routing alerts into maintenance systems so inspection, parts, labor, and technician disposition are captured.
- Alert Precision and Recall: Measures of whether alerts correspond to real faults and whether important faults are missed.
Pause at Figure 107.2 before carrying key concepts forward. Its visual vocabulary joins Three Maintenance Strategies to Average repair cost comparison across maintenance, which frames illustrative comparison of reactive, preventive, and predictive maintenance strategies.
Compare Three Maintenance Strategies with Average repair cost comparison across maintenance inside the visual at Figure 107.2. Next find $15K, which completes the scope of illustrative comparison of reactive, preventive, and predictive maintenance strategies. The decision in key concepts must preserve that labelled boundary.
Illustrative comparison of three maintenance strategies: reactive work waits for failure, preventive work follows a schedule, and predictive work uses condition evidence to schedule intervention before an expected fault.
This scenario timeline contrasts how the same equipment can behave under three maintenance regimes. The dollar values are illustrative inputs, not universal benchmarks.
Pause at Figure 107.3 before carrying equipment lifecycle comparison forward. Its visual vocabulary joins Maintenance Strategies Over 12 Months to Reactive, which frames timeline comparing three maintenance strategies across 12 months: reactive work carries outage risk, preventive work can replace healthy parts, and.
Compare Maintenance Strategies Over 12 Months with Reactive inside the visual at Figure 107.3. Next find Fix when broken, which completes the scope of timeline comparing three maintenance strategies across 12 months: reactive work carries outage risk, preventive work can replace healthy parts, and. The decision in equipment lifecycle comparison must preserve that labelled boundary.
107.9.1 Planning Cost Comparison
| Strategy | Example Cost Pattern | Budget Pattern | Unplanned Downtime Pattern |
|---|---|---|---|
| Reactive | Highest emergency repair exposure | Large unplanned share | Highest |
| Preventive | Scheduled replacement and inspection cost | Larger planned share | Lower, but parts may be changed early |
| Predictive | Sensor, analytics, and review workflow cost | Planned around condition evidence | Lower when alerts are trusted and actionable |
Interactive element unavailable — chart cell
Plot: Observable Plot (charting library) is not bundled
Show source
Plot.plot({
marginLeft: 100,
marginBottom: 60,
x: {
label: "Annual Cost ($)",
grid: true
},
y: {
label: null
},
marks: [
Plot.barX(comparison, {
y: "strategy",
x: "cost",
fill: d => d.strategy === "Reactive" ? "#E74C3C" : d.strategy === "Preventive" ? "#E67E22" : "#16A085",
tip: true,
title: d => `${d.strategy}\nCost: $${d.cost.toLocaleString()}\nSavings vs Reactive: $${d.savings.toLocaleString()}`
}),
Plot.text(comparison, {
y: "strategy",
x: "cost",
text: d => `$${(d.cost/1000).toFixed(0)}K`,
dx: -30,
fill: "white",
fontSize: 14,
fontWeight: "bold"
}),
Plot.ruleX([0])
],
color: {
legend: false
}
})Use the calculator above as an editable scenario model. For example, a plant might compare an assumed current maintenance factor with an assumed condition-based maintenance factor for a group of motors:
Current-state annual maintenance factor:
Condition-based annual maintenance factor:
Example maintenance-cost difference: $62,500 - $20,000 = $42,500
If the predictive maintenance system (sensors, installation, software, and integration) costs $180,000 upfront with $36,000 annual operating costs, this narrow maintenance-cost model produces:
The narrow maintenance-cost payback would be:
This does not prove the project is justified. It shows why the business case must include the local downtime rate, fault probability, spare-part lead time, false-positive cost, and whether the alert actually arrives early enough for planned work.
107.10 Predictive Maintenance Pipeline
The strategy comparison showed why condition evidence matters. The next question is how that evidence travels from a physical asset to a maintenance decision without losing context.
To test predictive maintenance pipeline, open the diagram in Figure 107.4. Predictive Maintenance Data Pipeline supplies one named condition; From sensor sampling to work orders and parts planning supplies the necessary comparison for predictive maintenance data pipeline from condition sensing to maintenance action.
-
Ada gathers condition evidence from the bearing.
-
The edge device turns the signal into useful features.
-
Analytics estimate fault type and severity.
-
The warning becomes a reviewable work order.
-
The technician result returns to the evidence loop.
Locate Predictive Maintenance Data Pipeline on Figure 107.4 before checking From sensor sampling to work orders and parts planning. The visual’s third anchor, 1. Sense, completes predictive maintenance data pipeline from condition sensing to maintenance action. Carry Predictive Maintenance Data Pipeline into predictive maintenance pipeline; use 1. Sense as its limiting condition.
Predictive maintenance data pipeline with four stages: condition sensors collect asset evidence, an edge device computes features, analytics estimate severity or remaining useful life, and a maintenance workflow turns the alert into reviewable work.
This diagram uses concrete example values to show where data volume changes. Treat the numbers as a design scenario that must be replaced by measured asset data in a real deployment.
Use Figure 107.5 to prepare the decision in alternative view: example data flow. The diagram names Predictive Maintenance Pipeline and Real data volumes from sensing through automated action, the two anchors needed to assess predictive maintenance pipeline with example data volumes: sensing captures raw condition signals, edge processing extracts features, analytics.
Use Real data volumes from sensing through automated action to test Predictive Maintenance Pipeline in the diagram at Figure 107.5. Then inspect Sensing as the final qualifier on predictive maintenance pipeline with example data volumes: sensing captures raw condition signals, edge processing extracts features, analytics. That sequence keeps alternative view: example data flow tied to what is visibly labelled.
107.11 Vibration Analysis
Rotating machinery (motors, pumps, fans) reveals health through vibration signatures:
The next claim about vibration analysis depends on Figure 107.6. Its diagram makes Vibration Analysis Workflow and From 3-Axis Accelerometer Sampling to Predictive explicit within vibration analysis workflow from sensing to defect detection.
Use From 3-Axis Accelerometer Sampling to Predictive to test Vibration Analysis Workflow in the diagram at Figure 107.6. Then inspect ACCELEROMETER as the final qualifier on vibration analysis workflow from sensing to defect detection. That sequence keeps vibration analysis tied to what is visibly labelled.
Vibration analysis workflow showing sensing (3-axis accelerometers at 100-1000 Hz) feeding both time-domain analysis (RMS, peak, crest factor) and frequency-domain analysis (FFT, order analysis, envelope analysis) to detect specific defects like imbalance, misalignment, and bearing faults.
107.11.1 Common Defects and Frequencies
| Defect | Frequency Signature | Detection Lead Time |
|---|---|---|
| Imbalance | 1x shaft speed | 1-2 weeks |
| Misalignment | 2x shaft speed (axial and radial) | Immediate |
| Bearing defects | BPFO, BPFI, BSF, FTF harmonics | 2-4 weeks |
| Gear mesh | Teeth count x shaft speed | 1-3 weeks |
| Looseness | Multiple harmonics, random spikes | 1-2 weeks |
107.11.2 Analysis Techniques
Time-domain analysis:
- RMS: Overall vibration level
- Peak: Maximum amplitude
- Crest factor: Peak-to-RMS ratio (indicates impulsive events)
Frequency-domain analysis:
- FFT: Fast Fourier Transform identifies specific defect frequencies
- Order analysis: Tracks frequency components relative to shaft speed
- Spectral trending: Monitors changes in specific frequency bands over time
Advanced techniques:
- Envelope analysis: Demodulates high frequencies to detect bearing faults
- Wavelet analysis: Time-frequency analysis for transient events
- Cepstrum analysis: Detects periodic patterns in spectrum (gear families)
107.11.3 Detection Timeline
Use this detection timeline section as a guided decision record, not as a list to memorise. First identify the stated input, assumption, or scenario; then compare each option on the same units and time boundary. Next check which value changes the outcome and which evidence would reveal an invalid assumption. For detection timeline, the useful result is the reasoning chain: observed condition, governing constraint, calculation or classification, and operational consequence. Record that chain before choosing an answer or carrying a value into the next section. Where the panel supplies several choices, reject each distractor against the chapter’s named mechanism instead of relying on wording cues. Where it supplies a table or timeline, compare rows at like-for-like scale and preserve the difference between an early indication, an actionable threshold, and a final outcome. This turns detection timeline into evidence that can be reviewed, recalculated, and connected to the running design narrative.
| Defect Type | Early Detection | Actionable Alert | Critical |
|---|---|---|---|
| Bearing wear | 6-8 weeks | 2-4 weeks | <1 week |
| Imbalance | 2-4 weeks | 1-2 weeks | Days |
| Misalignment | Immediate | Immediate | N/A |
| Lubrication | 4-6 weeks | 2-3 weeks | Days |
Interactive element unavailable — chart cell
Plot: Observable Plot (charting library) is not bundled
Show source
html`<div style="display: grid; grid-template-columns: 1fr 1fr; gap: 20px; margin: 20px 0;">
<div>
<h4 style="color: #2C3E50; margin-bottom: 10px;">Time Domain</h4>
${Plot.plot({
width: 400,
height: 250,
x: {label: "Time (seconds)", grid: true},
y: {label: "Amplitude", domain: [-2, 2]},
marks: [
Plot.line(vibrationSignal.slice(0, 200), {x: "time", y: "amplitude", stroke: "#16A085", strokeWidth: 1.5}),
Plot.ruleY([0])
]
})}
</div>
<div>
<h4 style="color: #2C3E50; margin-bottom: 10px;">Frequency Domain (FFT)</h4>
${Plot.plot({
width: 400,
height: 250,
x: {label: "Frequency (Hz)", grid: true},
y: {label: "Magnitude"},
marks: [
Plot.line(fftData, {x: "frequency", y: "magnitude", stroke: "#E67E22", strokeWidth: 1.5}),
Plot.ruleX([shaftSpeed], {stroke: "#2C3E50", strokeDasharray: "4 4"}),
Plot.text([{x: shaftSpeed, y: 0}], {
x: "x",
y: "y",
text: ["1x"],
dy: -10,
fill: "#2C3E50",
fontSize: 10
})
]
})}
</div>
</div>`The photographs below make infrared thermal-imaging camera a physical comparison: look for changes in package, exposed interfaces, mounting, scale, and service access before treating the forms as interchangeable.
Read across the forms as engineering evidence. They share a capability name, but packaging and installation change the electrical, mechanical, environmental, and maintenance constraints.
Checkpoint: Vibration Signals
You now know:
- Vibration analysis uses time-domain features such as RMS, peak, and crest factor plus frequency-domain methods such as FFT, order analysis, and envelope analysis.
- Common clues include 1x shaft speed for imbalance, 2x shaft speed for misalignment, and BPFO/BPFI bearing bands for race defects.
- Detection timelines differ: bearing wear may show early signs 6-8 weeks out, while misalignment can be immediate.
107.12 Continue to Part 2
Continue with Predictive Maintenance: Models and Rollout.
