A failure investigation that produces no written record is an investigation that produces no organisational value. The root cause identified by the engineer who spent three weeks on the problem disappears when that engineer changes roles, retires, or leaves the company. The corrective action that fixed the failure on product generation three is not available to the design team working on product generation six unless it was written down, stored somewhere findable, and actively reviewed before the new design was frozen.
This is the failure documentation problem, and it is the final discipline in the failure analysis chain — the step that closes the loop from field failure back to improved future design. The analytical methods covered in this series — visual inspection, failure mechanism identification, root cause analysis, FMEA, reliability engineering, and design for reliability — generate knowledge. Documentation and lessons-learned systems preserve that knowledge, make it accessible, and ensure it is applied. Without this final step, the value of the entire failure analysis process is degraded in proportion to how quickly the individuals who conducted the investigation move on from their roles.
Engineering organisations accumulate an enormous amount of failure knowledge through field service, warranty analysis, and product testing. Most of it is never systematically captured. It lives instead in email archives, in the notes of the engineer who worked on the problem, in the institutional memory of the warranty team, and in undocumented design decisions whose rationale becomes unclear within months of the project closing.
The consequences are predictable. The same failure mode appears on successive product generations because the connection between the field failure on the first generation and the design decision on the third was never explicitly made. A supplier change that triggered field failures on one product line is repeated on a different product line because the failure record was product-specific rather than supplier-specific. A material substitution that appeared safe based on room-temperature properties causes creep failures in service — because the investigation report from a similar substitution three years earlier was filed in a folder that no-one on the current design team knew existed.
The cost of each repeated failure includes not only the warranty and field service cost of the new failure, but also the sunk cost of the original investigation that produced knowledge that was then wasted. An organisation that re-learns the same lesson four times has paid four times the cost of discovering it once.
Institutional Memory Is Not a Reliability Strategy
The statement "ask Dave — he knows about the gearbox failures" is a symptom of
a documentation problem, not a solution to it. Institutional memory held by
specific individuals evaporates when those individuals change roles, retire,
or leave. Senior engineers who have been with a company for twenty years
accumulate decades of failure knowledge that is genuinely irreplaceable if it
is never written down. Failure documentation is not an administrative burden —
it is the mechanism by which individual experience becomes organisational
knowledge that outlasts the individual.
The failure report is the primary unit of failure documentation. A complete failure report contains five components.
Identification. The component or assembly that failed, the product it was installed in, the field application, the age at failure (operating hours, cycles, or calendar time), and the date of the investigation. This metadata enables retrieval by component type, application, and time period.
Failure description. What happened, from the user's perspective. What the product was doing when the failure occurred, what symptom was reported, and what the examination found. Photographs, dimensional measurements, and any forensic evidence collected. This section is factual and descriptive — no conclusions, no causes, no attribution.
Root cause. The specific mechanism by which the failure occurred and the specific reason that mechanism was present. The root cause must be stated at a level specific enough to point directly to a corrective action. "Material defect" is not a root cause. "PTFE gasket material specified without reference to the 380°C operating temperature, which exceeds the material's continuous service limit of 260°C" is a root cause. The root cause section is where the investigation analysis — the five whys, the fishbone, the Weibull parameters — produces its output in a single specific, attributable statement.
Corrective actions. The three levels: containment (what was done immediately to stop the failure affecting other units in service), corrective action (what was changed to prevent recurrence on the current design), and systemic action (what was changed in the process that allowed the root cause to exist — the specification procedure, the design checklist, the supplier qualification process). Each action must have a named owner and a completion date.
Verification. How the corrective action was confirmed effective — re-testing, follow-up field monitoring, Weibull life data after the design change, or documented absence of recurrence over a defined period. A report without a verification section is incomplete — it documents intention, not outcome.
A Failure Report Is Only as Good as Its Root Cause Section
A failure report that describes symptoms in detail but states the root cause
as "manufacturing defect" or "operator error" is a documentation exercise, not
an analysis record. The root cause section is the part that future design
teams need most — it tells them what specific condition to design against,
what parameter to verify, what specification to review. A vague root cause
produces a vague corrective action and a report that cannot be used to prevent
a similar failure on the next product. Require specificity: the root cause
must name the mechanism, the physical cause, and the systemic reason it was
allowed to exist.
The most technically rigorous failure report has no value if it cannot be found when it is needed. Retrievability is as important as content quality, and retrievability requires deliberate metadata design.
A failure report intended for retrieval must be tagged with at least four categories:
Component type: the category of component (seal, bearing, gear, fastener, electronic assembly, etc.)
Application or product family: the product line or application type in which the failure occurred
Root cause category: the system-level reason the failure occurred (design margin, material specification, manufacturing process, installation procedure, operating procedure, etc.)
These tags make it possible to ask the questions that matter for design decisions: "What have we seen with lip seals in outdoor hydraulic applications?" or "What failures have we had with this bearing supplier across all product lines?" or "How often has inadequate material temperature rating been a root cause?"
Without tags, a database of failure reports is a document archive — information is in it, but nobody can find it efficiently. With consistent tags applied at the time of filing, the same archive becomes a searchable knowledge base that actively informs future design decisions.
Tag Failure Reports for Retrieval, Not for Filing
The purpose of tagging is to make future retrieval fast and complete, not to
organise documents. Tag for the questions a future design engineer might ask:
what component, what mechanism, what application, what systemic cause. A
report tagged as "gearbox / wear / mobile equipment / inadequate lubrication
specification" will be found by the engineer designing a new mobile equipment
gearbox who searches for gearbox wear history. A report tagged only as
"product line X, 2023" is effectively unfindable outside a manual review of an
entire archive.
A lessons-learned database is a structured repository of failure reports, investigation findings, and corrective action records, designed for search and retrieval. The minimum viable structure for an engineering organisation managing complex products includes:
Full-text search across all report content — not just titles and tags. Failure reports that use specific technical language (seal extrusion, fretting corrosion, pitting fatigue) should be findable by any future engineer who searches those terms.
Faceted filtering by component type, failure mechanism, application, date range, and root cause category — using the metadata tags described above.
Cross-linking between failure reports that share a component type, supplier, failure mechanism, or root cause. A failure report on a gearbox lubricant issue should be linked to other lubricant-related failures across different products, not isolated within a single product line's records.
Status tracking for corrective actions — clearly indicating whether each action is pending, in progress, completed, or verified effective. An open corrective action is an unresolved risk; a database that does not track status cannot identify which investigations remain incomplete.
The database does not need to be sophisticated software. A well-structured shared spreadsheet with consistent column definitions and a controlled vocabulary for tags can serve the purpose for organisations with fewer than fifty failure reports per year. The critical requirement is consistent structure: every report in the same format, every tag from the same vocabulary, and a single point of entry so that all failure events are captured in one place rather than scattered across product-line folders and team drives.
Building the database solves only half the problem. The other half is ensuring that engineers actively consult it at the points in the design process where the knowledge is most valuable.
The highest-leverage integration point is the design review. For any design that uses components or materials from categories represented in the failure database, a mandatory lessons-learned review should be a gate before design freeze. The reviewer asks: what failures have we seen on similar components, in similar applications, with similar materials? The answer is often directly actionable — it surfaces minimum margin requirements, material exclusions, supplier alerts, and process controls that would otherwise only be discovered through field experience.
FMEA development is the second high-value integration point. An FMEA built without reference to the historical failure database starts from a blank sheet. An FMEA built by reviewing all relevant failure reports first starts from accumulated experience — the failure modes that have actually occurred are already in the analysis, and the occurrence ratings can be calibrated against historical failure rates rather than estimated from intuition alone.
New product introduction gating is the third integration point. Before a new product is released to production, a formal review of the lessons-learned database for similar products, similar failure mechanisms, and similar operating conditions identifies whether any known risks have been addressed or whether the new design may be repeating a historical mistake that was costly the first time.
The ultimate measure of a lessons-learned system is whether it changes future designs. A database that is consulted, but whose contents do not affect design decisions, is a searchable archive — not an active engineering memory.
The connection from failure report to design change operates through three mechanisms.
Design checklists. The recurring root causes in the failure database — inadequate material temperature rating, insufficient margin on fatigue-loaded fasteners, incorrect seal material for the service fluid — become line items on standard design checklists. Every engineer completing the checklist is prompted to verify the known-critical parameters. The checklist evolves as new failure reports add new root causes and existing ones are retired when designs make them obsolete.
Material and component exclusion lists. Specific materials, suppliers, or component types that have produced systematic failures are formally excluded from new designs through controlled lists referenced in the design standards. An exclusion list converts a failure history into a design constraint that cannot be ignored by accident or through ignorance of prior events.
Component-specific design requirements. For components with a documented failure history, design guidelines specify the minimum design margin, the material requirements, or the manufacturing controls required to achieve the target reliability. These guidelines are derived directly from failure investigation reports and the reliability analysis that followed — they represent the hardened institutional knowledge of every failure that category of component has experienced.
The Return on the Entire Series
The failure analysis methods in this series — visual inspection, failure
mechanism identification, root cause analysis, FMEA, reliability engineering,
and design for reliability — are investments. Their return is not the
individual failure corrected; it is the future failures prevented. Failure
documentation and lessons-learned systems are the mechanism that multiplies
the return on every investigation. Without them, each investigation saves one
product generation. With them, each investigation benefits every future
product generation that draws on the accumulated knowledge. The discipline
required is modest: write the report, tag it consistently, and review it
before design decisions are made.
A manufacturer of industrial drive systems experienced three separate gear failure events across successive product generations over eight years. Each failure was investigated: the first produced a recommendation to increase gear tooth surface hardness; the second produced a change to the lubricant specification after cold-temperature viscosity was identified as the cause of inadequate film formation; the third produced a design change to increase the gear centre distance and reduce contact stress at rated load. Each investigation was thorough and each corrective action was effective on the product generation where it was applied.
When the fourth generation product entered design, none of the engineers on the team had been involved in any of the three previous failure investigations. There was no structured failure documentation system. The reports existed — two in engineering investigation folders for old product lines, one in a warranty analysis spreadsheet in the quality team archive — but they were not indexed together, not tagged by failure mechanism, and not reviewed during the design process.
The fourth generation product repeated two of the three previous failure modes: inadequate surface hardness on the gear teeth and a lubricant viscosity specification that did not account for the minimum ambient operating temperature. Both failures were discovered in the field during the first year of service, triggering warranty campaigns on an installed base that had grown substantially since the first generation.
After the fourth generation failures, the organisation implemented a structured failure documentation system: a shared database with mandatory fields for component, mechanism, application, and systemic root cause; a policy requiring all investigation reports to be entered and tagged within thirty days of the investigation closing; and a design review gate requiring database review for any drive system product before design freeze.
The fifth generation product was designed against this database. The gear specification review surfaced the surface hardness requirement from the generation one failure report; the lubricant specification review flagged the cold-temperature viscosity requirement from generation two. Both were addressed before design freeze, at a cost of a few hours of engineering review time. No gear failures were recorded across the first two years of fifth generation field service. The cost of the documentation system — the template, the database structure, and the review gate — was approximately forty engineering hours to establish. The cost of the two fourth generation warranty campaigns it would have prevented was substantially higher.
Writing reports that describe symptoms but not root causes. A report that records that a part failed, documents the physical evidence, and concludes "insufficient material strength" provides no basis for preventing the next failure. The root cause must identify why the strength was insufficient — the wrong specification, wrong material supplied, unexpected service condition — at the specific level needed to design against it.
Treating the database as a filing system rather than a knowledge tool. A failure report filed and never retrieved has generated no organisational value. The value is in retrieval. If the database is never consulted at design review, something needs to change — either in how the database is structured and tagged or in how the design process gates access it.
Closing investigation reports before corrective actions are verified. A report that records that a corrective action was planned but does not record that it was effective is incomplete. The verification section — the field data, the test results, the documented absence of recurrence — is the evidence that the investigation produced lasting value. Without it, the report cannot confirm that the failure mode has been prevented.
Allowing failure reports to silo by product line. Failures on one product are informative for another product that shares common components, materials, or failure mechanisms. Cross-product retrieval requires a tagging and search structure that transcends product line boundaries. A database organised exclusively by product line rather than by failure mechanism will miss the cross-product patterns that represent the highest-value lessons — the systematic issues with a supplier, a material class, or a design practice that appear across multiple product families before anyone notices the pattern.
Failure documentation is the mechanism that converts expensive individual investigations into permanent organisational knowledge — without it, every failure investigation saves one product generation at most, and its value walks out the door with the engineers who worked on it
A complete failure report contains five components: identification, failure description, specific root cause, corrective actions at all three levels, and verification of effectiveness; reports without a specific root cause and verified corrective action are incomplete
Failure reports must be tagged at the time of filing with component type, failure mechanism, application, and root cause category — these are the search terms future engineers will use, and retrieval is as important as capture
The integration points that generate the most design value are design review gates, FMEA development seeded from failure history, and new product introduction reviews that explicitly check prior failure records for similar components and mechanisms
Design checklists, material exclusion lists, and component-specific design requirements derived from failure history are the mechanisms by which documented failure knowledge permanently changes how future products are designed