Last updated: 7 October 2026
Failure Mode and Effects Analysis (FMEA) helps facility and maintenance teams identify how an asset can fail, understand the consequence of each failure, and turn the highest risks into specific preventive maintenance inspections and work orders. For industrial and manufacturing facilities, FMEA creates a defensible basis for deciding what equipment to maintain first, what technicians should inspect, and when safety controls such as permits to work are required.
Key takeaways
- An FMEA failure mode is a specific way an asset, component, or process can fail, such as a pump bearing seizure or a compressed-air leak.
- Failure modes and effects analysis ranks risks using severity, occurrence, and detection scores, then focuses preventive maintenance on the most consequential and least detectable failures.
- A useful FMEA output is an actionable CMMS task: asset, inspection point, method, frequency, acceptance criteria, owner, and escalation path.
- Risk priority should guide work-order sequencing, while safety-critical work also requires appropriate isolation and e-Permit to Work controls.
- Teams should review FMEA scores after failures, inspections, modifications, and changes in production demand.
What is FMEA in facility maintenance?
FMEA is a structured method for identifying potential asset failures, evaluating their effects, finding their causes, and prioritising actions that reduce operational, safety, quality, or compliance risk. The full term is failure mode and effects analysis, although teams may also search for “failure effects mode analysis,” “failure mode FMEA,” or “FMEA analysis.” They refer to the same core risk-assessment discipline.
In a facility context, the asset under review may be a chilled-water pump, air compressor, boiler, conveyor, electrical switchboard, dust-collection system, or standby generator. The team breaks the asset into functional components, then asks four practical questions:
1. What function must this asset perform? 2. What is the potential failure mode? 3. What happens if that failure occurs? 4. What maintenance control can prevent, detect, or limit that failure?
For example, a cooling-water pump must maintain flow at a required pressure. A potential failure mode is bearing wear leading to excessive vibration. The effects may include reduced process cooling, unplanned production stoppage, seal damage, and motor overload. Planned vibration checks, lubrication verification, alignment inspections, and temperature monitoring become candidate maintenance tasks.
FMEA supports the asset-management principles of ISO 55000 by connecting asset risk to planned, repeatable decisions. It gives a maintenance programme a clearer rationale than applying the same calendar interval to every asset regardless of consequence.
How does FMEA differ from fault tree analysis?
FMEA examines potential failures from the bottom up, while fault tree analysis starts with an unwanted top-level event and traces the combinations of causes that could produce it. Both methods are valuable for industrial maintenance, and teams often use them together.
| Criteria | Failure Mode and Effects Analysis | Fault Tree Analysis |
|---|---|---|
| Starting point | A component, asset, or process function | A defined top event, such as total loss of compressed air |
| Direction of analysis | Bottom-up: failure mode to effect | Top-down: event to contributing causes |
| Main maintenance use | Creating inspections, PM tasks, and work-order priorities for many assets | Investigating a critical event and testing dependencies or combinations of failures |
| Typical output | Failure-mode register with scores, controls, and recommended actions | Logic diagram using AND/OR relationships between causes |
| Best fit | Building and improving an asset preventive-maintenance programme | Understanding why a high-consequence event can occur |
Use failure mode and effects analysis when a maintenance team needs a practical, repeatable way to establish PM work. Use fault tree analysis when a single event has serious safety, production, environmental, or business-continuity consequences and the team needs to understand interacting failures. For instance, an FMEA can identify generator battery degradation as a failure mode; a fault tree can explore every condition that could result in generator failure during a utility outage.
What are the core steps in a facility asset FMEA?
A facility asset FMEA should move from asset function to controlled maintenance actions through a consistent seven-step workflow. This approach, called the FMEA-to-PM Action Loop, ensures that risk scoring leads to work that technicians can perform and managers can audit.
1. Define the asset boundary and required function
The FMEA team should define the equipment included, operating context, interfaces, and performance requirement before listing failures. “Air compressor” is too broad if the receiver, dryer, filters, controls, and distribution header have different functions and risks.
Record the asset ID, location, manufacturer, duty or standby role, operating hours, upstream and downstream dependencies, and required output. A compressor supplying control air to process valves has a much different consequence profile from one serving workshop tools.
2. Identify each failure mode
A failure mode describes the specific way the asset fails to meet its required function. Good entries are observable and technically clear: “motor fails to start,” “pump delivers insufficient flow,” “filter differential pressure exceeds limit,” or “control panel trips intermittently.”
Avoid vague entries such as “equipment failure.” That wording cannot become a useful inspection checklist or work-order instruction. The phrase failure mode effective analysis sometimes appears in searches, but effective FMEA requires precise modes, effects, causes, and controls rather than broad labels.
3. State the local and wider effects
The effect explains what the failure means for the component, asset, process, people, and site. Teams should document both immediate and downstream effects. A failed belt on an extraction fan may stop airflow locally, increase dust exposure, disrupt production, create a fire risk, and trigger environmental non-conformance.
Write effects in operational language that production, EHS, and maintenance leaders can validate. Clear effects make severity scoring more consistent.
4. Identify causes and current controls
The FMEA team should identify credible causes, such as contamination, misalignment, loose terminals, lubrication breakdown, corrosion, operator error, software configuration, blocked vents, or ageing insulation. Then record existing preventive controls and detection methods.
A control may be a scheduled visual inspection, thermal scan, oil analysis, vibration route, alarm, interlock test, or operator round. If a control exists only as tribal knowledge or an unstructured spreadsheet reminder, it is difficult to prove completion or assess whether findings received follow-up.
5. Score severity, occurrence, and detection
The team should score severity (S), occurrence (O), and detection (D) using one agreed scale, commonly 1 to 10. Severity measures the consequence of the effect. Occurrence estimates how likely the cause is. Detection evaluates how likely the current controls are to detect the issue before the effect reaches the operation.
The traditional Risk Priority Number is calculated as:
`RPN = Severity × Occurrence × Detection`
An RPN helps sort a long failure register, yet it should never override safety or regulatory judgment. A low-frequency failure with severe injury potential deserves escalation even if its numerical RPN is lower than a frequent nuisance fault. Many teams therefore set a separate rule: any high-severity item requires action, review, or formal acceptance by an accountable leader.
6. Select a maintenance response
The FMEA team should choose the control that best addresses the failure mechanism. Possible responses include an inspection, condition-monitoring route, functional test, replacement interval, operator care task, engineering modification, spare-parts decision, training action, or contingency procedure.
The response must match the cause. Replacing a filter on a fixed interval may control contamination, while thermal imaging may better reveal a developing loose electrical connection. A failure mode that cannot be cost-effectively prevented may require a planned corrective strategy, a critical spare, and a rapid-response work-order process.
7. Review after evidence changes
An FMEA is a living record that should be updated when a fault occurs, a PM inspection finds deterioration, a design changes, or production conditions shift. Each completed work order supplies evidence about whether occurrence, detection, task frequency, or task content should change.
How do you turn FMEA findings into preventive maintenance tasks?
FMEA findings become useful when each selected control is converted into a clear, recurring task in a CMMS. The work order must tell a technician what to inspect, how to judge the result, and what happens when the result falls outside limits.
A robust preventive-maintenance task includes:
- Asset and location: Link the task to the correct equipment hierarchy and physical area.
- Failure mode addressed: State the risk, such as “bearing overheating from inadequate lubrication.”
- Inspection or maintenance instruction: Use a defined action, such as checking grease condition, vibration level, guard condition, and bearing temperature.
- Method and acceptance criteria: Specify the approved tool, measurement point, safe operating condition, tolerance, or pass/fail standard.
- Frequency and trigger: Set a calendar interval, runtime interval, condition trigger, or production-cycle trigger based on risk and evidence.
- Assigned competency: Route electrical, mechanical, instrumentation, or contractor work to qualified personnel.
- Evidence requirement: Require readings, photos, meter values, checklists, or completion notes where appropriate.
- Escalation rule: Generate corrective work when an inspection fails, and define the priority and approver.
A digital inspection checklist makes this process operational. Technicians can complete standard steps at the asset, flag defects, attach evidence, and raise follow-up work without re-entering the issue in another system. Facility teams building this foundation can use preventive maintenance schedules for facility assets to establish workable intervals, asset records, and review routines.
How should a maintenance team prioritise FMEA work orders?
A maintenance team should prioritise FMEA-derived work orders by consequence, loss of control, and operational window, rather than by the age of the request alone. The aim is to prevent the highest-impact failure modes before they affect safe, reliable production.
| Work-order priority | FMEA signals | Typical action |
|---|---|---|
| Emergency | Immediate safety danger, environmental release, active critical-service loss, or mandatory shutdown condition | Make safe, isolate where required, dispatch authorised responders, then document recovery and root cause |
| Urgent | High severity with failed or absent detection control; deteriorating condition likely to create a near-term outage | Plan the earliest safe production window, reserve parts, and assign qualified technicians |
| Planned priority | Moderate-to-high RPN with effective temporary control; task can be completed within a scheduled maintenance window | Issue a planned corrective work order and verify completion before the next review date |
| Routine PM | Lower consequence or well-controlled failure mode; recurring inspection is the primary control | Execute through the PM schedule and escalate only when acceptance criteria fail |
Safety requirements sit alongside work-order priority. Electrical isolation, confined-space entry, hot work, line breaking, or work around hazardous energy may require a permit process before a technician begins. A digital permit workflow provides approvals, precautions, isolation records, and closure evidence. For a deeper guide, see permit to work software for safer, auditable facilities.
Which facility assets should receive FMEA first?
Facility teams should start FMEA with assets whose failure can stop production, create harm, damage high-value equipment, breach compliance obligations, or cause major recovery effort. Starting with the entire asset register usually delays action and creates an overly generic analysis.
For an industrial or manufacturing site, the initial shortlist often includes:
- Main electrical incomers, switchgear, transformers, UPS systems, and emergency generators
- Production-critical pumps, motors, conveyors, hydraulic units, and process-control equipment
- Boilers, chillers, cooling towers, compressed-air systems, and process ventilation
- Fire detection, suppression, emergency lighting, extraction, and life-safety systems
- Dust collectors, wastewater treatment equipment, chemical handling systems, and environmental controls
- Assets with repeated reactive work orders, expensive downtime, or single-point-of-failure status
Use maintenance history, operator feedback, fault reports, production dependencies, and EHS records to validate the shortlist. A CMMS makes this easier when asset history, recurring PMs, inspection outcomes, and corrective work orders are linked to the same equipment record.
How can a CMMS support FMEA-based maintenance?
A CMMS supports FMEA-based maintenance by connecting risk decisions to assigned, scheduled, traceable work. The FMEA worksheet itself is only an analysis artefact; the operational value arrives when recommended controls become work orders and their results improve future scoring.
FacilityBot’s CMMS, preventive-maintenance schedules, digitised checklists, and work-order management can support this workflow by recording assets, assigning recurring inspections, capturing findings, and routing defects into accountable corrective work. Messaging-based fault reporting can also bring technician or operator observations into the maintenance record quickly, where supervisors can assess whether a reported symptom maps to a known failure mode.
When evaluating software, test whether the platform can preserve the relationship between the asset, failure mode, task, result, corrective action, and completion evidence. Buying decisions should also account for configuration scope, asset-data readiness, integrations, and user roles. Review FacilityBot pricing alongside the workflow requirements and rollout plan, rather than considering subscription cost in isolation.
How often should facility teams review an FMEA?
Facility teams should review high-risk FMEAs after any relevant failure, major inspection finding, process change, or maintenance strategy change, with a scheduled periodic review for all active analyses. The correct interval depends on asset criticality and how quickly operating conditions change.
Review meetings should compare predicted risks with work-order evidence. If an inspection continually finds no degradation, the team may adjust frequency after technical review. If defects repeatedly appear between planned visits, the team may need a shorter interval, a different detection method, improved access, a redesign, or a condition-monitoring approach.
The review should also confirm that technicians can execute the specified task safely and consistently. A well-written FMEA that produces impractical inspections will not improve reliability.
FAQ
What is an example of an FMEA failure mode for a facility asset?
An FMEA failure mode for a chilled-water pump could be “pump delivers insufficient flow due to impeller wear.” Its effects may include inadequate cooling, process temperature excursions, higher energy use, and production interruption. Suitable controls could include flow trending, vibration monitoring, differential-pressure checks, and planned impeller inspection during shutdown.
Is FMEA preventive or predictive maintenance?
FMEA is a risk-analysis method that can support both preventive and predictive maintenance. It may result in time-based inspections and replacements, or it may identify condition-monitoring methods such as vibration analysis, thermography, oil analysis, and sensor alarms as the most effective detection controls.
What should happen when an FMEA inspection fails?
When an FMEA inspection fails, the technician should record the finding against the asset, create or trigger a corrective work order, apply the defined priority, and make the asset safe if the condition presents an immediate risk. The maintenance planner should later use the result to review the original FMEA score and control effectiveness.
A practical FMEA programme turns maintenance risk into visible, completed action. To map your asset risks into digital inspections, preventive schedules, and accountable work orders, book a FacilityBot demo.