Every engineering component will eventually fail in some way. The question is not whether failure modes exist but which ones are worth preventing, in what order, and at what cost. Failure modes and effects analysis (FMEA) is the structured method for answering those questions before the product is built and before customers experience the consequences.
FMEA is a proactive tool — it applies the analytical discipline of failure analysis not to a part that has already failed but to a design or process that has not yet been built or released. It asks systematically: for each component, what could go wrong? What are the consequences if it does? How likely is it to occur? Can it be detected before it causes harm? The answers to these questions, ranked by risk, produce an action list that directs design and process improvements to the highest-priority failure modes.
FMEA is a risk prioritisation tool. It identifies and ranks potential failure modes by their combined risk — the product of how severely they would harm the customer, how frequently they are likely to occur, and how detectable they are before causing harm. The output is a ranked list of failure modes with corresponding recommended actions.
FMEA is not a guarantee against failure. A completed FMEA does not mean a product will be reliable — it means the team has systematically identified and ranked the known failure modes and taken action on the highest-priority ones. Unknown failure modes are not in the FMEA. Novel designs, novel materials, and novel use conditions produce failure modes that were not anticipated in the original analysis. FMEA reduces the risk from known and foreseeable failure modes; field experience and reliability testing reduce the risk from unexpected ones.
FMEA is also not a pass/fail test. There is no threshold score above which a product is "safe" and below which it is not. The tool is a ranking and prioritisation mechanism, and its value lies in what changes it drives, not in the final numerical scores it produces.
FMEA Is a Team Activity
A useful FMEA is not written by one engineer in isolation. It requires input
from design engineers (who understand how the product is intended to work),
manufacturing engineers (who understand the process variations that affect
quality), quality engineers (who understand historical failure modes from
similar products), and service technicians (who understand how the product is
actually used and maintained in the field). A single-author FMEA consistently
misses failure modes known to other team members. Plan for two to four hours
of structured team review for each subsystem.
A structured FMEA follows a consistent sequence of steps applied to each component or process step in scope.
Define the scope and level of analysis. FMEA can be performed at the system level (interactions between major assemblies), the subsystem level (interactions between components in an assembly), or the component level (failure modes of individual parts). Start with the level where design decisions are still being made — system-level FMEA too early produces vague results; component-level FMEA too late catches failures after the design window for major changes has closed.
Identify the function of each component or process step. A failure mode is a deviation from the intended function. Without a clear function statement, it is impossible to systematically identify failure modes. Write the function as a verb-noun statement: "transmit torque from motor shaft to gearbox input," "seal fluid at 10 bar pressure," "locate bearing inner ring axially."
Identify potential failure modes for each function. A failure mode is the way in which the component could fail to perform its function: shaft fractures, seal leaks, bearing inner ring migrates axially. List all plausible failure modes, not just the most obvious ones. Review field failure data, warranty records, and failure data from similar designs for this step.
Identify the effects of each failure mode. The effect is what the customer or downstream system experiences when the failure mode occurs. Effects are written from the customer's perspective: "machine stops unexpectedly," "fluid leak creates safety hazard," "excessive vibration at output shaft." A single failure mode may have multiple effects at different levels — local effect (what happens at the component), subsystem effect, and system effect (what the customer experiences).
Identify potential causes of each failure mode. The cause is the mechanism by which the failure mode could occur: "shaft fractures due to fatigue from stress concentration at keyway root," "seal leaks due to extrusion damage during installation," "bearing inner ring migrates due to insufficient interference fit at operating temperature." Specific, root-cause-level causes produce actionable FMEA; generic causes like "material defect" or "manufacturing error" produce nothing useful.
Three numerical ratings are assigned to each cause-failure mode-effect combination. Each is rated on a scale of 1 to 10.
Severity (S) rates the consequence to the customer if the failure mode occurs and reaches the customer. 10 = safety hazard or non-compliance with regulations. 9 = loss of primary function with potential safety implication. 7–8 = significant performance reduction or customer very dissatisfied. 5–6 = reduced performance, some customer dissatisfaction. 3–4 = minor inconvenience. 1–2 = negligible effect, customer may not notice.
Severity is fixed by the failure mode and its effects — it does not change based on controls or improvements to detection. If the failure mode has severity 9, it is rated 9 regardless of how good the detection controls are.
Occurrence (O) rates the probability that the specific cause will occur over the design life. 10 = failure is almost certain (>1 in 2 units affected). 8–9 = high likelihood. 6–7 = moderate likelihood. 4–5 = occasional occurrence. 2–3 = low probability. 1 = extremely unlikely.
Occurrence should be based on data wherever possible — field failure rates from similar products, process capability data (Cpk), or accelerated testing results. Estimated occurrence ratings without data are less reliable but still useful for relative ranking.
Detection (D) rates the ability of current controls to detect the failure mode or its cause before it reaches the customer. 10 = no controls exist; the failure mode will certainly reach the customer undetected. 7–9 = low likelihood of detection. 4–6 = moderate likelihood. 2–3 = high likelihood of detection. 1 = almost certain detection by existing controls.
Low Detection Score Does Not Make a Failure Mode Safe
A failure mode with severity 9, occurrence 2, and detection 1 has an RPN of 18
— the same as a failure mode with severity 3, occurrence 3, and detection 2.
The RPN treats them equally, but they are not equally concerning. Any failure
mode with severity 8 or above should receive immediate attention regardless of
occurrence and detection scores. Detection controls reduce the probability
that a failure reaches the customer; they do not change the severity of the
consequences when it does.
The Risk Priority Number (RPN) is the product of severity, occurrence, and detection ratings:
RPN = S × O × D
The maximum RPN is 10 × 10 × 10 = 1,000. In practice, most FMEA entries have RPNs in the range of 10 to 400.
The RPN provides a relative ranking of failure modes within the FMEA to direct attention and resources. High-RPN items should receive recommended actions to reduce one or more of S, O, or D. The goal is not to reduce every RPN to zero (which is impossible) but to ensure that the highest-risk failure modes have appropriate controls or design changes.
Reducing severity: redesign to eliminate the failure mode or change the effect. If the failure mode cannot be eliminated, design fail-safe features that reduce the effect severity. Severity reductions are the most valuable because they reduce the worst-case consequence.
Reducing occurrence: redesign to make the failure mode less likely. Use more robust materials, tighter manufacturing tolerances, more conservative design margins. Reference designs that have demonstrated low failure rates in similar service.
Reducing detection: add inspection, testing, or monitoring controls that identify the failure mode or its cause before the product reaches the customer. In-process controls detect problems during manufacturing; design verification tests detect problems before production release; field monitoring (sensors, alerts) detects problems in service before they become safety issues.
Two main FMEA types address different stages of product development.
Design FMEA (DFMEA) analyses potential failures in the product design — how the product could fail to perform its intended function due to design decisions. DFMEA is conducted during the design phase, before the design is frozen, so that recommended actions result in design changes. The scope is the product design and the intended operating conditions.
Process FMEA (PFMEA) analyses potential failures in the manufacturing or assembly process — how the process could produce a non-conforming product. PFMEA is conducted during process planning, before production tooling is committed, so that recommended actions result in process controls, mistake-proofing (poka-yoke), or inspection steps. The scope is the manufacturing process steps and the product characteristics they affect.
A complete quality plan for a new product includes both DFMEA (starting at concept or preliminary design) and PFMEA (starting at detailed process design). The DFMEA informs the PFMEA by identifying which product characteristics are most critical to prevent design failures — these characteristics become the priority control items in the PFMEA.
The difference between a useful FMEA and a documentation exercise is in the recommended actions and their implementation.
Every failure mode with an RPN above a defined threshold — or with severity 9 or 10 regardless of RPN — must have a recommended action that is specific, assigned to an owner, and has a completion date. Vague recommended actions ("improve design," "add inspection") are not acceptable — they cannot be tracked or verified.
Recommended actions should address root causes, not symptoms. An action that adds a final inspection step addresses detection; it does not change the likelihood that the failure mode occurs. Actions that redesign to eliminate the failure mode, increase the design margin, or mistake-proof the process address occurrence and are more valuable.
The FMEA must be updated after recommended actions are implemented and re-evaluated. The revised severity, occurrence, and detection ratings produce a revised RPN that confirms whether the action was effective. An FMEA that is written, approved, and then never updated does not reflect the current design and loses its value as a living risk management tool.
Link the FMEA to the Test Plan
The detection controls in a DFMEA should directly inform the design
verification test plan. If the FMEA identifies a failure mode with a high
detection score (hard to detect in design verification), the test plan should
include a specific test designed to detect that failure mode. An FMEA and a
test plan that were developed independently will have gaps — test cases that
do not align with the identified failure modes and failure modes with no
corresponding test coverage.
A pneumatic rotary actuator was being designed for an outdoor valve automation application. A DFMEA was conducted on the pneumatic circuit and rotary mechanism. One of the failure modes identified during the team review was as follows.
Component: piston rod seal. Function: retain air pressure and prevent moisture ingress. Failure mode: seal leaks. Effect: actuator loses torque under load, valve position cannot be maintained (severity 7 — valve could fail to close in an emergency). Cause A: seal extrusion during assembly due to inadequate chamfer on housing bore entry. Cause B: seal degradation from UV and ozone exposure in outdoor service.
For Cause A (assembly extrusion): Occurrence rated 5 (occurs occasionally without assembly tooling); Detection rated 6 (final pressure test does not always detect small extrusion damage that worsens over time). RPN = 7 × 5 × 6 = 210.
Recommended action for Cause A: add a 15-degree lead-in chamfer with a minimum length specified on the drawing, and add a mandrel assembly tool to the assembly work instruction. Expected revised occurrence: 2 (requires tool misuse to damage seal). Revised RPN = 7 × 2 × 6 = 84.
For Cause B (UV/ozone degradation): Occurrence rated 4 (outdoor exposure over five-year design life); Detection rated 7 (no current detection of seal degradation without disassembly). RPN = 7 × 4 × 7 = 196.
Recommended action for Cause B: change seal material from NBR (poor ozone resistance) to EPDM or silicone (excellent UV/ozone resistance for outdoor service). Add seal condition to the periodic maintenance inspection checklist. Expected revised occurrence: 2. Revised RPN = 7 × 2 × 5 = 70.
Both recommended actions were implemented before design freeze. No seal failures were reported in the first two years of field service across forty-two installed units.
Rating severity based on current controls. Severity reflects the inherent consequence of the failure mode reaching the customer. It does not change based on how good the detection controls are. Lowering a severity rating because "we have a good inspection process" is incorrect and dangerous — it understates the risk if the detection control fails.
Treating RPN as the only criterion for action. As noted above, severity 9–10 failure modes require action regardless of RPN. A failure mode that will certainly be detected (detection = 1) and rarely occurs (occurrence = 2) but would cause a safety hazard (severity = 9) has an RPN of 18 — but it must be addressed.
Writing causes at too high a level. "Manufacturing defect" is not a cause. "Shaft diameter undersized due to tool wear not detected by current sampling frequency" is a cause. Specific causes produce specific recommended actions; generic causes produce generic and ineffective responses.
Closing the FMEA before verifying recommended actions. An FMEA with recommended actions that have not been implemented or verified does not reflect the actual risk level of the design. The FMEA must be updated — with revised ratings and action completion status — after each recommended action is completed.
The FMEA Is Worth What It Changes
An FMEA that results in no design changes, no process improvements, and no new
test cases has value only as documentation. An FMEA that results in three
design changes, two added inspection controls, and five additional test cases
has real engineering value — it has measurably reduced the risk of field
failures that would otherwise cost far more to address after product release.
The return on FMEA investment is proportional to the quality and
implementation rate of the recommended actions, not to the completeness of the
documentation.
FMEA predicts failure modes proactively, before they occur. When failure does occur in the field — despite FMEA, despite design validation, despite manufacturing controls — root cause analysis is the tool for understanding why and preventing recurrence. The next post covers root cause analysis methods: the five whys and the fishbone (Ishikawa) diagram, and how to apply them to reach the systemic root cause rather than stopping at the proximate symptom.
FMEA systematically identifies and ranks potential failure modes before they occur, driving design and process improvements to the highest-risk items
Risk Priority Number (RPN = Severity × Occurrence × Detection) provides relative ranking, but severity 9–10 items require action regardless of RPN
Severity reflects the inherent consequence of the failure mode; it does not change based on detection controls
Design FMEA (DFMEA) addresses product design failures; Process FMEA (PFMEA) addresses manufacturing process failures — both are needed for a complete quality plan
An FMEA that results in no design changes, no process improvements, and no additional test coverage has documentation value only — the return on investment comes from implemented corrective actions