The start of 2026 marked 20 years since my graduation from my Masters degree in Reliability Engineering. At that stage “reliability engineer” was an uncommon term in Australian industry, but I’m pleased to say that it has grown much more common over time and now almost every asset-intensive organisation has at least one position with that title. I’m sure that the inclusion of “Reliability Engineering” in the early editions of the GFMAM’s Asset Management Landscape was a large factor, so thanks to the asset management visionaries that led that work.
While “brand recognition” has grown, it seems obvious to me that Reliability Engineering is still in its infancy in general industry. In most organisations it is treated as a role rather than a discipline – and one that is often filled by an engineer with very limited knowledge of the concepts and principles that underpin sound Reliability Engineering practice. Core techniques – from reliability modelling to root cause analysis – are poorly applied or not applied at all. This leaves money on the table for most organisations – and sometimes even exposes them to significant safety or environmental risks. Accordingly, I’ve planned out a series of articles to address the most common and significant issues that I see. I’m calling this “Reliability Engineering Excellence” and I’m starting it with a view of reliability across the life cycle.
Reliability Engineering is More than Failure Reduction
Before we go any further, I want to clarify my scope a little bit. Without getting into definitions from standards, “reliability” relates to the probability of an asset failing, while “Reliability Engineering” is a broader concept, covering all sorts of techniques used to manage the consequences of asset failure, usually by either preventing, mitigating, or accepting those consequences. This expands the focus of Reliability Engineering to include Reliability, Availability, Maintainability, and even Supportability (you may have heard “RAM” or “RAMS”) as characteristics of an asset that can be studied and manipulated to manage failure. Safety and risk are fully integrated into this view as integral constraints rather than optional extras. Trying to solve (or even understand) asset performance challenges with a narrow reliability focus is fighting with one arm tied behind your back.
In taking a broad view of Reliability Engineering there is clearly overlap with various other professions and disciplines, including Systems Engineering, Risk Engineering/Management, Maintenance Engineering, and, of course, Asset Management . Individual organisations need to define their roles, responsibilities, and operating models to match their context, so feel free to mentally allocate the activities I discuss to other roles. Just make sure that those roles have access to the rigorous probability tools and models that are the reliability engineer’s stock in trade.
One more point before we get to the details. When it comes to Reliability Engineering, pragmatism matters. There are some very complex analytical tools and methods available, and these are valuable, but they do not need to be applied in every situation. The right approach depends on the organisation – its purpose, objectives, history, current situation, and future expectations. A simple reliability block diagram (RBD) may be enough to expose the production constraint. A targeted RCM or PM optimisation review may be more valuable than a full strategy rebuild. A defect elimination process may deliver the fastest benefit. A sophisticated model is only useful if it changes a decision. Accordingly, I will focus this series on the most common tools that add real, practical value to typical organisations. Your needs may differ and I’d be happy to discuss them.
A Lifecycle View of Reliability Engineering
Most organisations are focussing their Reliability Engineering efforts on the maintenance phase, but the truth is that RAMS outcomes are created, protected, and sometimes eroded across the entire lifecycle. Different tools apply in different phases, so let’s step through each in turn.
Reliability Engineering in Concept & Acquisition
The concept and acquisition phases are particularly critical for RAMS because this is the most cost-effective time to incorporate changes that influence the rest of the lifecycle. As shown in the following diagram, most of the lifecycle cost (which is usually dominated by maintenance costs to prevent and correct failure) is locked in by early design decisions.

Key activities for reliability engineers during this phase are:
- Specification – All design activities should include RAMS requirements as part of the specification. It’s not enough to guess – the requirements need to focus on what constitutes “success”, which might be dominated by either availability or reliability, with the other RAMS factors playing a role in optimising for this outcome.
- Modelling – Tools like RBDs can cover all elements of RAMS and are useful to drive top level requirements down into sub-systems and components, as well as to estimate in-service performance. Failure Mode Effects and Criticality Analysis (FMECA) is a well-known (but under-utilised) tool to design out failure modes and, for those that are retained, to drive maintenance decisions using Reliability Centred Maintenance (RCM) or Risk Based Inspection (RBI) principles. These tools – and a huge range of others – allow designers to identify, explore, and resolve RAMS risks before the equipment enters service. For example, the very simple RBD below already tells you that the reliability of the system will be driven by the lowest reliability component and design improvements should focus there.

- Design Principles – Adoption of dedicated design principles can assist designers to address the risks and achieve the requirements. These include:
- Redundancy, derating, and failure tolerance to improve reliability
- Accessibility, visibility, testability, complexity, interchangeability, identification/labelling, verification, and simplicity to improve maintainability
- Simulation – You could reasonably consider simulation a form of modelling, but I’ve broken it out
- Testing – Testing for RAMS performance is rarely done due to perceived expense, but could well be the key to avoiding a costly design mistake. This goes well beyond concepts like “wet commissioning” and requires a statistically valid test period with pre-defined acceptance criteria. Expense can be reduced through Accelerated Life Testing, focussing on high-risk components, and Bayesian methods.
Reliability Engineering in Operations
Some organisations have their reliability team reporting to the operations manager, emphasising the importance of Reliability Engineering in continuity of production. Even if they’re not part of operations, the reliability team needs to be actively monitoring the key operating parameters to ensure they remain within expected limits as these are key to achieving the expected RAMS performance. Change in operations is fine, but it needs to be controlled so that the right maintenance, support equipment, spares, and even replacement projects are in place – and so that management can balance the costs of these against the benefits of the change.
In addition to monitoring and reworking the RAMS performance models, reliability engineers can assist operations by monitoring observed failures and feeding back where operators could or should have acted differently. A classic example is shown below, where the handle on a valve has been allowed to corrode entirely away. Regardless of whether this is a case of not reporting damage or not releasing equipment for maintenance, it didn’t happen in a few weeks or even a few months and now represents a serious issue if/when that valve needs to be operated. As the saying goes – schedule your maintenance, or it will schedule itself for you.

There are a few specific initiatives/approaches that are often led by the reliability team and can be useful in driving good behaviours in the operations team. Two of the most common are:
- Operator Driven Reliability (ODR) – an approach where operators are actively involved in routine inspections and minor maintenance, taking advantage of their familiarity with the equipment and improving their ownership.
- Tighten, Lubricate, Clean (TLC) – playing off “tender loving care”, encourages both operators and maintainers to practice basic asset hygiene that can be highly effective in preventing failures.
An excellent example of the value to be gained is illustrated below. In this work, a relatively minor reduction in dump truck load variability was achieved by proper training and monitoring of loader operators and truck drivers. This work simultaneously increased the average load per trip and reduced the overload rate, increasing productivity and reducing failure rates.

Reliability Engineering in Maintenance
Regardless of where they are sitting, the reliability team need to be actively involved in maintenance. As with operations, this starts with monitoring high level parameters (e.g. availability, overall equipment efficiency) and extends from there to identifying and resolving performance issues with the assets. Most commonly, these will arise where the failure rate, type, or consequence deviates from expectations and will need investigation using Root Cause Analysis (RCA) techniques to determine the causes and solutions. RCA needs to be wrapped in a larger process – often referred to as Defect Elimination (DE) – that ensures that required changes are implemented in a controlled manner, with the operator behaviour changes discussed above being just one example.
Reliability engineers should generally facilitate the RCA process and, where necessary, can support this and broader business improvement activities with specific skills such as:
- Weibull analysis, degradation analysis, or other curve fitting techniques to bring statistical rigour to critical decision making.
- Reliability Centred Maintenance (RCM) or Preventive Maintenance Optimisation (PMO) to correct any shortcomings in the maintenance program. This is, of course, based on Nolan & Heap’s famous study of the frequency of the six failure patterns in maintenance and concluded that age-based replacement is generally a poor choice, as illustrated below.

- Spares analysis to optimise inventory holding costs against downtime and other costs from stock outs
- Design tools to revisit any of the design decisions where there’s a gap between expected and observed performance
Quite a few of these skills involve the probability and statistics tools that I discussed above, but there’s also potential to bring value with expertise around those tools. For example, good reliability engineers will be aware of Resnikoff’s Conundrum, which notes that we are less likely to have good information on more critical failures because we’ve already done quite a lot to prevent them. That’s a good outcome, but it does mean that we always have to be vigilant about how we apply statistical information in practice.
Reliability Engineering in Replacement and Disposal
Surprising as it may seem, the lead-up to replacement or disposal can be an excellent opportunity for the reliability team to deliver value. Firstly, there is the repair or replace decision, attempting to determine if the asset needs to be replaced and, if so, when. This is usually an economic analysis of the aggregated productivity, failure and cost information, or occasionally a detailed analysis of a few critical failure modes. The broad options are:
- Extend life – requiring analysis of risk, with treatment through operational limits or additional maintenance
- Overhaul – requiring detailed analysis of failure modes
- Upgrade – requiring a smaller scale design activity, with RAM requirements
- Replace – requiring aggregation of RAM data to inform the requirements for the replacement
As can be seen, the reliability engineers will be well placed to assist with this decision making.
Even when the choice is simply to dispose of the asset, reliability engineers can assist. I recently worked with a team that were within months of shutdown of their plant, and it was remarkable just how many reliability engineering decisions they were having to revisit every day. This was essential to ensure that they maintained safety and took sensible production risks as they worked to reduce expenditure in line with the limited remaining life of the assets. A handful of examples:
- When can specific maintenance activities be stopped?
- Should a spare that has just been issued be re-ordered?
- Does the maintenance workforce need to be increased or decreased as the maintenance program moves from preventive to reactive?
How Mature is Your Organisation? Six Questions to Ask…
If you’re wondering whether your organisation is making good use of your reliability engineering capability, then try the following six questions. If you can confidently answer “yes” to all of these then you are in a good position – and definitely well beyond the maturity of most organisations I have worked with.
- Do you know what RAM performance requirements your assets were designed to deliver?
- Do you know what your current RAM performance is?
- Do you know whether you are experiencing unexpected types or rates of failure?
- Do your operators use your assets the way you expect them to?
- Do your maintainers maintain your assets the way you expect them to?
- Have you adjusted your operations and maintenance programs to reflect upgrades, modifications or life extensions?
If the answer to one or more of the above is “no” then look out for some detailed and practical guidance on how to address this in the rest of this series. I’m going to start with a discussion about the reliability improvement journey, which I’m calling “It’s Never Too Late to Get Better”.
