Root Cause Analysis: Five Whys and Fishbone Method
Root Cause Analysis: Five Whys and Fishbone Method
Joshua R. Lehman
Author
Failure Analysis15 min read
When a component fails in the field, the first instinct is to replace it and return the equipment to service. This response addresses the symptom. It does not address why the component failed, and it does nothing to prevent the next failure — which will occur on the same schedule as the first, under the same conditions, for the same reason. Root cause analysis breaks this cycle by identifying the systemic cause behind the failure event so that corrective action targets the mechanism, not just the outcome.
Root cause analysis is not an investigation into what happened — the physical failure mechanisms covered in earlier posts address that question. RCA is an investigation into why it happened in a way that allows the failure to be prevented from happening again. The distinction matters because a failure that is understood physically but not systemically will recur.
Every failure has a proximate cause — the immediate physical mechanism that produced the damage — and one or more root causes — the systemic conditions that allowed the proximate cause to develop. The proximate cause is what you see on the fracture surface or worn component. The root cause is what created the conditions for that damage to occur.
Consider a hydraulic cylinder that developed seal leakage after eighteen months of service. The proximate cause is seal wear — the seal lip is worn and no longer maintains contact pressure. But this tells you nothing useful for prevention. Why was the seal worn? Because the fluid was contaminated with particulate. Why was the fluid contaminated? Because the reservoir breather filter had not been replaced at the specified interval. Why had the breather filter not been replaced? Because it was not on the maintenance checklist for that equipment. Why was it not on the checklist? Because the checklist was developed from a generic template that did not account for the specific accessories on this model.
The proximate cause is seal wear. The root cause is a maintenance checklist that was not tailored to the equipment. Replacing the seal without correcting the checklist produces the same failure at the same interval. Correcting the checklist prevents the next failure.
Root Cause Is Not Always Singular
Most engineering failures have multiple contributing causes, not a single root
cause. The five whys and fishbone methods are both designed to surface
multiple causal pathways. A thorough RCA identifies all contributing causes
and evaluates each one for corrective action — not just the most obvious or
the easiest to fix. When a corrective action addresses only one of three
contributing causes, the failure rate may decrease but will not be eliminated.
The five whys method is an iterative interrogation technique developed at Toyota as part of the Toyota Production System. It works by repeatedly asking "why" in response to each cause statement until the analysis reaches a systemic root cause — one where the corrective action would prevent the failure from recurring, not merely repair the immediate damage.
The method is simple in structure: state the problem, ask why it occurred, state the cause, ask why that occurred, and continue until the answer is a systemic failure — a process gap, a specification deficiency, a design assumption that was never validated, a training or communication failure — rather than a proximate physical mechanism.
Problem statement: Hydraulic cylinder seal failed after eighteen months of service.
Why 1: Why did the seal fail? The seal was worn beyond serviceable limits. The lip contact force was insufficient to maintain the seal.
Why 2: Why was the seal worn? The hydraulic fluid was contaminated with particulate that abraded the seal lip.
Why 3: Why was the fluid contaminated? The reservoir breather filter was blocked and had been passed its replacement interval by approximately six months.
Why 4: Why was the breather filter not replaced? The breather filter was not listed on the preventive maintenance schedule for this equipment.
Why 5: Why was it not on the maintenance schedule? The maintenance schedule was copied from a base template and not reviewed against the equipment's accessory list before being issued.
Root cause: the procedure for generating equipment-specific maintenance schedules does not include a step to verify the schedule against the actual accessory and consumable list for each machine variant.
Corrective action: revise the maintenance schedule generation procedure to require a line-by-line review against the equipment bill of materials. Update existing schedules for the affected equipment family. Add a review step to the new equipment commissioning checklist.
Five Iterations Is a Guide, Not a Rule
The name "five whys" suggests the analysis always requires exactly five
iterations. In practice, some failures reach a systemic root cause in three
iterations; others require seven or eight. The correct stopping point is when
the answer is a systemic failure — a broken process, a missing specification,
a training gap — rather than a physical mechanism or a proximate condition.
Stopping too early leaves the causal chain incomplete. Stopping arbitrarily at
five when the root cause has not been reached produces a corrective action
that targets a symptom.
The five whys method is powerful for problems that have a clear linear causal chain — one condition caused another, which caused another. It works well for process failures, maintenance deficiencies, and quality escapes where the sequence of events is straightforward.
It is less effective for complex failures where multiple independent causes converged. A product that failed because of a material deficiency, a stress concentration that was not in the design specification, and an operating condition that exceeded the design basis cannot be adequately described by a single linear chain of whys. Following any one causal strand to its root cause misses the contribution of the other two. The fishbone diagram addresses this limitation.
Five whys also depends heavily on the quality of each answer. If any link in the causal chain is wrong or incomplete, the entire analysis goes astray. The technique requires honest, specific answers at each step — not speculation, not blame assignment, and not stopping at the first answer that feels complete.
The fishbone diagram, also called the Ishikawa diagram after its developer Kaoru Ishikawa, organises potential causes of a failure into a structured visual framework. The effect (the failure event) is written on the right side of the diagram. Major causal categories extend as branches from a central spine, and specific causes are listed as sub-branches under each category.
The most common causal categories for engineering failures are the six M's: Man (human factors — training, procedure following, decision-making), Machine (equipment and tooling condition), Method (process and procedure), Material (raw material and purchased component properties), Measurement (inspection and testing adequacy), and Environment (temperature, humidity, contamination, operating conditions).
Not all six categories are relevant to every failure, and the categories can be adapted to the specific context. For a manufacturing quality problem, Machine and Method are often the richest categories. For a field failure with an environmental dimension, Environment and Material may be primary.
The value of the fishbone diagram over the five whys is breadth. By systematically populating each causal category, the team surfaces potential causes that would not emerge from a single linear interrogation. A fishbone session run with representatives from design, manufacturing, quality, and service typically identifies twenty to forty potential causes — far more than a single-author five whys analysis.
Populate the Fishbone Before Evaluating Causes
The most common mistake in fishbone analysis is evaluating each cause as it is
generated — deciding immediately whether it is relevant, plausible, or worth
investigating. This stops the idea generation prematurely. The correct
technique is to populate all branches of the diagram first, generating as many
potential causes as possible without evaluating them, and then to evaluate and
rank the causes as a second step. The brainstorming phase and the evaluation
phase should be kept separate.
The output of the fishbone session is a populated diagram with all potential causes identified. The next step is to evaluate each cause for likelihood and evidence. Each cause is assessed against available evidence — physical inspection findings, measurement data, operating records — and ranked by probability. The highest-probability causes become the focus of root cause verification: testing, examination, or data analysis to confirm or refute the hypothesis.
The two methods are complementary and are most effective when used together. The fishbone identifies the range of potential causes; the five whys drills down each confirmed cause to its systemic root.
A typical combined approach proceeds in four steps. First, use the fishbone diagram in a team session to brainstorm all potential causes, categorised by the six M's. Second, evaluate each potential cause against the available evidence and identify the two or three most likely root causes. Third, apply the five whys to each confirmed cause to trace it to its systemic root. Fourth, develop corrective actions targeted at the systemic roots identified.
This approach produces broader, more complete RCA than either method alone. The fishbone prevents the five whys from following only one causal strand when multiple contribute. The five whys prevents the fishbone from stopping at proximate causes without reaching the systemic level.
Root cause analysis produces value only through the corrective actions it generates. A completed fishbone diagram with no corresponding actions is documentation, not problem-solving. Corrective actions from RCA must meet four criteria to be effective.
Address the root cause, not the symptom. An action that replaces the worn seal addresses the proximate cause. An action that adds the breather filter to the maintenance schedule addresses the root cause. Only the second prevents recurrence.
Be specific and verifiable. "Improve maintenance procedures" is not a corrective action. "Add breather filter replacement (part number X) at 1,000-hour intervals to the maintenance schedule for equipment models A, B, and C, effective date Y, verified by maintenance manager sign-off" is a corrective action. The distinction determines whether the action can be implemented, tracked, and confirmed complete.
Have an owner and a completion date. An action without an assigned owner will not be completed. An action without a date will be completed when it is convenient, which means never.
Be verified for effectiveness. After implementation, the corrective action should be evaluated to confirm that the failure mode has been eliminated or its frequency reduced. An action that addresses the correct root cause will produce a measurable change in failure rate. An action that does not reduce the failure rate indicates either that the root cause identification was incomplete or that implementation was inadequate.
Corrective Actions Have Three Levels
The hierarchy of corrective action matches the permanence of the solution.
Level 1 — containment — stops the immediate damage but does not prevent
recurrence (replace the failed part, quarantine affected product). Level 2 —
corrective action — addresses the root cause to prevent recurrence of this
specific failure mode (update the maintenance schedule). Level 3 — systemic
corrective action — addresses the systemic process gap that allowed the root
cause to exist (revise the new equipment commissioning checklist so that no
maintenance schedule is issued without a bill-of-materials review). All three
levels may be required. Level 3 is the most valuable because it prevents the
same root cause from producing different failure modes in other equipment.
A hydraulic power unit on a mobile work platform was experiencing repeated hydraulic pump failures, with five pump replacements required over fourteen months of service. Each pump was replaced as a warranty claim, and each was found to have internal scoring consistent with inadequate lubrication. The operating fluid was the specified grade, changed at the specified interval.
Step 1 — Fishbone diagram construction:
A team session with field service, design engineering, and the pump supplier produced the following candidate causes under the six M's:
Man: technician installing incorrect pump model; incorrect fill procedure leaving air in the system; incorrect fluid type used during one service event
Machine: pump operating outside design pressure range; hydraulic reservoir level consistently low; return line restriction elevating pump case drain pressure
Method: commissioning procedure not specifying pump prime-out steps; maintenance log not capturing fluid condition at change
Material: hydraulic fluid not meeting the pump supplier's viscosity requirement at minimum ambient temperature
Measurement: no fluid condition monitoring; no pressure gauge in the pump case drain line
Environment: mobile platform operating in an ambient temperature range extending to −25°C in winter
Step 2 — Evaluation against evidence:
Physical examination of the failed pumps consistently showed scoring on the port plate and cylinder block at locations consistent with metal-to-metal contact during startup. The damage pattern was localised to the earliest wear marks on the port plate face, consistent with inadequate fluid film during startup rather than sustained running contamination.
Operating records showed that four of the five pump failures occurred within the first week of service following a scheduled maintenance event. This clustering around post-maintenance startups was not previously noted.
Fluid viscosity data showed that the specified fluid (ISO VG 46) had a kinematic viscosity of approximately 430 cSt at −25°C — the minimum winter ambient. The pump supplier's specification required a maximum startup viscosity of 860 cSt, which the fluid met. However, the pump required ten to fifteen minutes of low-load operation to reach operating temperature before full-load use was permissible. The commissioning procedure and the maintenance return-to-service procedure made no mention of this warm-up requirement.
Step 3 — Five whys on the confirmed cause:
Why 1: Why did the pump fail? Inadequate lubrication film during high-load operation at startup.
Why 2: Why was lubrication inadequate? Fluid was too viscous at low ambient temperature to maintain hydrodynamic film at full load during cold startup.
Why 3: Why was full load applied before the fluid reached operating temperature? No warm-up procedure was in place; operators applied load immediately after startup.
Why 4: Why was there no warm-up procedure? The pump supplier's installation manual specified a warm-up requirement, but it was not incorporated into the equipment commissioning procedure or the maintenance return-to-service procedure.
Why 5: Why was the supplier requirement not incorporated? The equipment commissioning procedure was written from an internal template, and the review process did not include a check of supplier installation manuals for operating restrictions.
Root cause: the procedure for developing and reviewing equipment commissioning and maintenance procedures did not include a step to verify compliance with supplier installation requirements for key components.
Step 4 — Corrective actions:
Level 1: Replace the current failed pump. Implement a manual warm-up requirement in the interim operating instruction, distributed to all field operators immediately.
Level 2: Revise the commissioning procedure and the maintenance return-to-service procedure to include a ten-minute low-load warm-up cycle before full-load operation when ambient temperature is below +5°C.
Level 3: Revise the procedure development process to require a review step verifying that critical supplier installation restrictions (viscosity limits, warm-up requirements, startup procedures) are captured in the equipment operating procedures before they are issued.
After implementing the revised warm-up procedure, no further pump failures occurred over the following twenty-two months of service across fourteen units.
Stopping at the proximate cause. "The pump failed because the fluid was too viscous" is a proximate cause, not a root cause. A corrective action that changes the fluid grade without addressing why the warm-up requirement was not known or followed will produce a different failure mode at the next cold weather season.
Assigning blame rather than analysing systems. "The operator did not warm up the pump" is not a root cause — it is a symptom of a training, procedure, or communication failure. RCA that terminates at human error without asking why the human error was possible has not reached the systemic root. Systems that rely on operators remembering undocumented requirements fail when people are busy, new, or distracted.
Working alone. A single-author RCA consistently misses causes known to other team members. Service technicians know what operators actually do in the field. Supplier representatives know the limits of their components. Design engineers know what assumptions were made. An RCA conducted without input from all relevant parties is systematically incomplete.
Generating corrective actions that cannot be verified. "Improve training" is not a corrective action. "Add pump warm-up requirements to the operator training module, deliver to all current operators by date X, and add to onboarding for new operators" is a corrective action. The difference determines whether the action produces a measurable change in failure rate.
The Goal Is Prevention, Not Documentation
Root cause analysis has one purpose: preventing recurrence. An RCA report that
is filed without implementing corrective actions, or that implements
corrective actions without verifying their effectiveness, is documentation
with no engineering value. The measure of a successful RCA is not the quality
of the causal diagram or the depth of the report — it is whether the failure
recurs. Track corrective action implementation rates and failure recurrence
rates for each RCA. If recurrence is common, the RCA process is stopping too
early or producing corrective actions that are not implemented.
Root cause analysis addresses specific failures after they occur. The next post moves from reactive to proactive: reliability engineering provides the quantitative tools for predicting how long a population of components will survive, identifying the failure rate characteristics of a design, and making evidence-based decisions about maintenance intervals, warranty periods, and design life targets. The key tools are mean time between failures (MTBF), the Weibull distribution for modelling life data, and the methods for analysing field failure populations to extract reliability parameters.
Root cause analysis distinguishes the proximate cause (what failed physically) from the root cause (the systemic condition that allowed the failure) — corrective action must target the root cause to prevent recurrence
The five whys method follows a causal chain iteratively until it reaches a systemic failure such as a process gap, specification deficiency, or training omission — not a proximate physical mechanism
The fishbone diagram structures potential causes across six categories (Man, Machine, Method, Material, Measurement, Environment) and prevents the analysis from following only one causal strand when multiple contribute
Combining the two methods — fishbone to identify the range of causes, five whys to drill each confirmed cause to its systemic root — produces more complete RCA than either method alone
Corrective actions must be specific, assigned, dated, and verified for effectiveness — vague actions and unverified implementations produce no reduction in failure recurrence rate