Design for Reliability: Building Robustness into Products
Design for Reliability: Building Robustness into Products
Joshua R. Lehman
Author
Failure Analysis13 min read
Reliability is not a property that can be inspected into a product or tested into it at the end of the design process. By the time a product reaches design validation testing, the geometry is fixed, the materials are specified, the manufacturing process is defined, and the major design decisions that govern service life have already been made. Testing at this stage confirms what the design achieves — it does not change what the design achieves. The design decisions that determine reliability must be made earlier, during the concept and detailed design phases, using tools that translate reliability targets into quantifiable design parameters.
Design for reliability (DfR) is the collection of methods that makes this proactive approach possible. It applies the knowledge accumulated through failure analysis — the failure mechanisms, their causes, and their sensitivities to design parameters — to new designs before those designs have accumulated any field experience. The result is a product whose reliability is predictable, verifiable, and aligned with its target before the first unit reaches a customer.
The series up to this point has covered the mechanisms by which engineering products fail: fatigue, corrosion, wear, overload, thermal degradation. It has covered the methods for identifying why failures occur and preventing recurrence: FMEA, root cause analysis, Weibull life data analysis. Design for reliability is the integration of these tools into the forward-looking design process.
The connection is direct. An FMEA conducted during design identifies the failure modes most likely to cause field failures and their risk priority. Weibull parameters from similar products indicate the expected failure rate shape — whether wear-out is likely and at what characteristic life. Stress-strength analysis quantifies the design margin against each critical failure mode. Accelerated life testing validates the design before committing to production tooling. The tools are not independent methods — they are a connected system that transforms field failure knowledge into design robustness.
Reliability Must Be Specified Before It Can Be Designed
The first step in design for reliability is specifying the reliability target
in quantitative terms: R(t) — the required survival probability at a specified
service life, under specified operating conditions. Without a quantitative
target, there is no basis for evaluating whether a proposed design meets the
requirement. Qualitative reliability goals ("the product should be reliable")
are not design targets. Quantitative targets ("R(8,000 hours) ≥ 0.95 under
rated operating conditions") are testable, verifiable, and can be directly
linked to design margin requirements.
The fundamental model of mechanical reliability is the stress-strength interference model. Every component has a strength distribution — the distribution of strengths across a population of nominally identical parts, varying due to material property scatter, dimensional variation, and surface condition variation. Every component experiences a stress distribution — the range of operating stresses across a population of identical applications, varying due to load variation, environmental variation, and usage pattern variation.
Failure occurs when the applied stress on a specific unit exceeds the strength of that unit. When the stress and strength distributions are well separated — the mean strength is much higher than the mean stress, and the distributions do not overlap — the probability of failure is low. When the distributions overlap, the overlap area is proportional to the probability of failure.
This model reveals what a nominal safety factor does not show: two designs can have the same ratio of mean strength to mean stress but completely different reliabilities if their distributions have different widths. A design where both strength and stress are tightly controlled (low variance) can achieve high reliability with a safety factor of 2.0. A design where either distribution is wide (high variance in strength or stress) may require a safety factor of 3.0 or higher to achieve the same reliability.
To reduce the probability of stress-strength interference:
Increase the mean design margin (raise mean strength or lower mean stress)
Reduce the strength variance (tighter material specifications, tighter manufacturing tolerances, surface finishing to reduce variability)
Reduce the stress variance (better load characterisation, load limiters, design features that distribute load more uniformly)
The Tails of the Distributions Govern Reliability
The probability of failure in a stress-strength model is determined by the
overlap of the distribution tails, not the overlap of the distribution bodies.
A Gaussian distribution with mean strength of 200 MPa and standard deviation
of 10 MPa has negligible probability below 170 MPa — three standard deviations
from the mean. If the maximum expected stress is 180 MPa, the nominal safety
factor is 1.11 — seemingly inadequate. But if the stress distribution is
narrow (standard deviation of 5 MPa), the overlap is small and reliability is
high. Characterising the tails of both distributions — through sufficient
material testing and load measurement — is essential for accurate reliability
prediction. Assuming Gaussian distributions when the tails may be non-Gaussian
(heavy-tailed materials, rare overload events) is a common source of
reliability prediction error.
The design safety factor is the ratio of the design strength to the design load, where "design strength" is a lower bound on the strength distribution and "design load" is an upper bound on the stress distribution. A safety factor of 1.5 means the design strength is 50 percent above the design load — providing a buffer against the inevitable variability in both.
Safety factors in established design standards encode the accumulated experience of the industry with variability in materials, manufacturing, loading, and use conditions. ASME pressure vessel standards, structural steel codes, and lifting equipment standards each specify minimum safety factors calibrated to the typical uncertainty in those fields. Applying safety factors below the code minimums without thorough statistical characterisation of the specific stress and strength distributions is not conservative engineering — it is unknown territory.
The appropriate safety factor depends on four factors. First, the consequence of failure — a safety-critical failure justifies a higher factor than a performance failure. Second, the uncertainty in the load — a well-characterised, controlled load requires a smaller buffer than a variable, poorly characterised load. Third, the uncertainty in the strength — a well-characterised material with tight property control requires a smaller buffer than a new material or a variable process. Fourth, the failure mode — ductile yielding failure modes are more forgiving than brittle fracture or fatigue, because ductile materials redistribute stress locally before failure.
Derating is the practice of specifying a component for operation at a fraction of its rated capacity, intentionally leaving a margin between the operating condition and the rated limit. It is standard practice in electronic component reliability and is increasingly used in mechanical design.
A hydraulic cylinder rated to 350 bar maximum operating pressure derated to a maximum system pressure of 280 bar — 80 percent of rating — operates with a 25 percent pressure margin that reduces seal and port stress, decreases fatigue loading on the cylinder barrel, and extends the service life of wear items. The cost of the derating is the need for a larger cylinder to produce the same force; the benefit is a measurably longer service life and higher reliability.
Derating is most effective for components whose failure rate is sensitive to a key parameter — pressure, temperature, current density, contact stress. Reducing that parameter by 20 percent may extend service life by a factor of two or more, depending on the stress-life relationship for the dominant failure mechanism.
Derate Against the Dominant Failure Mechanism
Derating is only effective if it targets the parameter that drives the
dominant failure mechanism. A bearing derated by reducing radial load while
operating at excessive temperature receives no benefit from the load reduction
if thermal degradation of the lubricant is the life-limiting mechanism.
Identify the dominant failure mechanism first (from FMEA or historical failure
data), then derate the parameter that most strongly drives that mechanism.
Blanket derating without mechanism identification may miss the critical
parameter entirely.
Accelerated life testing (ALT) applies elevated stress to a sample of units to produce failures faster than would occur under normal service conditions. The failures provide life data for Weibull analysis; the Weibull parameters are then extrapolated to service conditions using a stress-life model to predict field reliability.
Three types of stress acceleration are used most commonly in mechanical product testing.
Elevated load or pressure. For components with a power-law stress-life relationship (fatigue, wear, surface fatigue), increasing the applied load accelerates damage accumulation. The acceleration factor — the ratio of test life to field life — is calculated from the ratio of test stress to service stress raised to the stress-life exponent. For fatigue with an S-N slope of b = 0.1 (typical for steel), increasing stress by 40 percent reduces the life by a factor of approximately (1.4)^(1/0.1) ≈ 1.4^10 ≈ 29. A test that produces failures in 200 hours at elevated stress predicts 5,800 field hours at rated stress.
Elevated temperature. For failure mechanisms driven by chemical kinetics — polymer degradation, seal material hardening, lubricant oxidation — the Arrhenius model relates acceleration factor to temperature: AF = exp[Ea/k × (1/T_use − 1/T_test)], where Ea is the activation energy, k is the Boltzmann constant, and T are temperatures in Kelvin. A typical activation energy of 0.7 eV produces an acceleration factor of approximately 10 for a 20°C temperature increase.
Combined stress. Real products experience multiple simultaneous stress factors. Combined ALT applies two or more elevated stresses simultaneously, using a combined acceleration model. The interactions between stress factors must be understood — some combine multiplicatively (independent mechanisms), others interact synergistically (corrosion and fatigue combine to produce corrosion fatigue at a rate greater than either alone).
Reliability growth testing is a structured test-analyse-fix (TAAF) cycle applied iteratively to identify and eliminate failure modes before release. A prototype or pre-production sample is tested under representative service conditions, often with some accelerated loading, until failures occur. Each failure is root-cause analysed, a corrective action is implemented, and the test continues with the corrective action incorporated. The reliability target is not met when the test starts — it is grown toward the target through iterative improvement.
The reliability growth model (Crow-AMSAA model) tracks cumulative failures against cumulative test time on a log-log plot. A straight line on this plot indicates a consistent improvement rate. The slope of the line, combined with the target reliability, determines when the test should stop — when the extrapolated reliability reaches the target.
Reliability growth testing is resource-intensive but produces designs with directly measured, improving reliability backed by test evidence. It is standard practice in military and aerospace procurement because the cost of field failures in those applications dwarfs the cost of pre-release testing. For industrial products with high warranty exposure or safety implications, the same discipline applies.
A manufacturer of industrial valve actuators was developing a new electric linear actuator for outdoor process plant service. The target reliability was R(20,000 hours) ≥ 0.90 under rated conditions (full rated thrust at maximum ambient temperature of +55°C, minimum ambient −30°C, IP67 enclosure).
The FMEA identified four high-priority failure modes: motor winding insulation degradation at high temperature, leadscrew wear in contaminated environments, output seal extrusion on high-force cycles, and circuit board condensation corrosion under temperature cycling.
For the motor winding insulation (Arrhenius mechanism, activation energy 1.0 eV), an accelerated thermal aging test was designed at 130°C. At the service temperature of 55°C, the acceleration factor was AF = exp[1.0/8.617×10⁻⁵ × (1/328 − 1/403)] ≈ 125. To accumulate the equivalent of 20,000 service hours, the test required 20,000 / 125 = 160 hours at 130°C. The ten-unit test produced two failures at 140 and 165 equivalent service hours. Weibull analysis of these two failures with eight suspensions gave β = 3.1, η = 28,000 equivalent service hours. R(20,000 hours) = exp[−(20,000/28,000)^3.1] ≈ 0.88 — marginally below target.
The corrective action was to upgrade the winding insulation class from Class F (155°C) to Class H (180°C), which raised the rated temperature capability and produced a higher effective activation energy. The retest produced no failures at the same test duration, with a revised lower-confidence-bound reliability estimate exceeding 0.93 at 20,000 hours.
This single test loop — test, analyse, correct, retest — identified a marginal insulation specification and corrected it before production tooling was committed. The alternative, discovering insulation failures in the field after three years of service, would have required a field retrofit campaign at substantially greater cost.
Designing to a nominal safety factor without considering variability. A safety factor of 2.0 applied to a material with a wide strength distribution and a loading environment with high variability may produce a reliability no better than 0.80. Quantifying the distributions and calculating the resulting reliability — rather than relying on a nominal factor — reveals whether the margin is adequate.
Treating accelerated test results as pass-fail rather than life data. A test that runs to a scheduled duration without failure produces a lower confidence bound on reliability but no positive information about the shape of the failure distribution. Running tests to failure, even with small samples, produces Weibull parameters that enable life prediction. Zero-failure tests are less informative per unit of test time than tests designed to produce a controlled number of failures.
Fixing observed test failures without addressing the failure mode family. A reliability growth test that corrects each specific failure without asking whether the root cause applies to other components produces a product that passes the test but fails in the field for related reasons. Each corrective action should be evaluated for applicability across the design — a seal extrusion failure on one joint may indicate a design practice that should be reviewed at all joints.
Reliability Is an Investment, Not a Cost
The cost of design for reliability — FMEA time, stress-strength analysis,
accelerated life testing — is typically two to five percent of the product
development budget. The cost of inadequate reliability in the field — warranty
claims, field service, recall campaigns, reputation damage, and in
safety-critical applications, liability — routinely exceeds the development
budget for a poorly reliability-engineered product. Companies that treat DfR
as a cost to be minimised discover its value only when the bill arrives in
warranty claims. Companies that treat it as an investment discover it when
their warranty costs are a fraction of their competitors'.
Design for reliability uses quantitative tools to build robustness into products before they leave the design phase. The final post in this series closes the loop from failure back to the organisation: failure documentation and lessons-learned systems ensure that every field failure, every root cause analysis, and every corrective action generates permanent knowledge that improves future designs. The value of the entire failure analysis discipline depends on this final step — turning individual failure events into organisational learning that prevents the same failure from occurring in the next product generation.
Design for reliability translates quantitative reliability targets into specific design decisions during the concept and detail design phases — not after design validation testing reveals deficiencies
Stress-strength analysis quantifies the overlap between the stress and strength distributions; reliability depends on the tails of both distributions, not just their means
Safety factors must account for variability in both stress and strength — the same nominal safety factor produces very different reliabilities depending on the width of the distributions
Derating reduces the operating stress below rated capacity, extending life for components whose failure rate is sensitive to a key parameter; it must target the parameter driving the dominant failure mechanism
Accelerated life testing produces Weibull parameters that predict field reliability; reliability growth testing iteratively eliminates failure modes through test-analyse-fix cycles before production commitment