In the first article in this series, I described Reliability Engineering as a lifecycle discipline and emphasised that the greatest opportunity to influence reliability often exists during asset acquisition. That’s a useful framework, but the practical reality is that most organisations are not starting with a clean sheet of paper. They have existing assets, existing maintenance strategies, existing operating practices, existing data quality issues, existing organisational habits, and a long list of competing priorities. If that sounds familiar, this article is for you. I’m going to discuss some tools, tips, and tricks that can get you started with meaningful RAM (Reliability, Availability, and Maintainability) improvements when you’ve already got in-service assets. The first question on that journey has to be – do you even have a RAM problem?
Do You Have a RAM Problem?
For all the strengths of Reliability Engineering, it is still just one discipline and it can’t solve every business problem. That’s why it’s important to start here. If you don’t, but attempt to push through a reliability improvement initiative anyway, even an obvious “good practice” initiative like introducing Reliability Centered Maintenance (RCM) for maintenance strategy decision making, you’re likely to run into opposition from others. They will want you to justify the investment against other priorities and you may struggle to do so, possibly even harming the reputation of Reliability Engineering within the organisation. I know because I’ve been guilty of this myself and experienced exactly that outcome – a great process, reluctantly approved, with no resources to implement it, resulting in no change. It’s frustrating and it doesn’t help anyone.
Avoiding this pitfall starts with a strong alignment between your organisation’s objectives and the improvement initiatives being proposed. This is a topic we discuss regularly at Assetivity – including a detailed discussion from an asset management perspective in our Asset Management Value Roadmap series. As I discussed in the first article in this series, Reliability, Availability and Maintainability (RAM) objectives should not sit separately from the organisation’s asset management objectives. They are a critical part of them.
Once that alignment is established, it becomes much easier to identify where assets are underperforming against business requirements, which is the foundation for a robust and defensible case for change. Without that connection, reliability initiatives can easily be seen as technical improvements in search of a problem. With it, they become business improvement initiatives supported by reliability engineering.
Work the Problem Before Choosing the Solution
Even when you’ve established a gap between the RAM performance, you’re getting and the RAM performance you want, it’s still not time to jump straight into a solution. I’ve seen countless organisations move directly from recognising poor performance to selecting an improvement methodology. They decide they need an RCM. Or predictive analytics. Or lean maintenance. Sometimes they’re right; often they’re not. It’s worth investing time upfront to understand the true causes of the problem so that you have confidence in the solution you’ve selected and the evidence to demonstrate why it is appropriate.
In practice, this often looks like a Root Cause Analysis (RCA) exercise. One of my preferred starting points is a simple Ishikawa, or fishbone, diagram. While these diagrams are commonly used to investigate specific technical failures, they’re equally effective when exploring business problems and organisational capability issues. Here’s an example I prepared in Copilot for a hypothetical haul truck availability problem:

The example illustrates how these tools support systematic exploration of the factors that may be contributing to poor reliability performance. You can then look for evidence to determine which potential causes actually exist, then investigate their deeper root causes. For example, clear evidence of systematic overloading of trucks would raise questions about operator training, load monitoring systems, supervision, and feedback. If the evidence points to operator practices and beliefs as a significant contributor, then an Operator Driven Reliability program could be an appropriate response. It’s a case of following the evidence, rather than your assumptions.
Using Data and Statistics
As you explore your RAM problem, you may find it necessary to do some detailed analysis of failures or other data – and this is another way the reliability team should be helping. In this case, it’s just a matter of applying the appropriate statistical or modelling skills. I was going to say, “to produce the required outputs”, but be very careful about that kind of statement – fixing the problem means getting to the truth, not reinforcing our preconceived opinions! This means using statistical tools and data analysis in the right way, always being aware of the assumptions and limitations of the techniques.
As the very first step in that process, you should never let your reliability engineers (or any other data analysts) state their results as a single number – e.g. “The MTBF of the system is 1,232 hours”. They always need to include an indicator of certainty such as error bars or confidence intervals that clearly expose the limitations of the technique and the data. As I like to say, “All models are wrong, some models are useful” and a model that does not provide an indication of certainty is very hard to make useful. Reliability Engineering doesn’t eliminate uncertainty. It helps organisations understand and manage it.
Choose the Intervention That Fits the Opportunity
Once you understand the problem and its causes, selecting an appropriate intervention becomes much easier. Different reliability challenges require different reliability tools, so I can’t cover them all in this article – or even in this series – but here are some of the more common issues I see and my recommended approaches to solving them.
Equipment Capacity Is Unclear
Many organisations monitor performance against nameplate capacity but have little understanding of what their assets can realistically achieve under current operating, maintenance and support conditions. They invest in maintenance improvements, additional spares, and external support without first understanding the capability of the overall system, resulting in high costs with limited improvement.
When I encounter this situation, my preferred approach is to revisit or often create Availability Models and Reliability Block Diagrams. These models help determine what the equipment is actually capable of achieving under current conditions. If the analysis shows performance cannot achieve business requirements, the answer is probably capital investment rather than further maintenance optimisation.
Maintenance Strategies That Don’t Prevent Failures
Incorrect maintenance strategies generally arise for one of two reasons: either the right strategy was never selected in the first place, or changes in operating conditions, asset age, workforce capability or asset configuration have altered failure behaviour over time. Regardless of the cause, the correct approach is to revise the maintenance strategies.
In most cases, a targeted Preventive Maintenance Optimisation (PMO) process is a more practical and effective solution than a traditional RCM review. By focusing on the failure modes that matter most, organisations can improve results more quickly while also identifying underlying changes in asset behaviour or asset utilisation.
It’s also worth reviewing broader capabilities such as operational readiness, configuration management, change management, and defect elimination. After all, understanding how the maintenance strategy became misaligned can be just as valuable as correcting it.
Growing Inventory with Declining Service Levels
I’ve seen many organisations face the uncomfortable combination of rising inventory values and declining material availability. Often this occurs because inventory practices have not kept pace with changes in asset configuration or maintenance requirements. The result is obsolete stock, poor inventory decisions, and declining demand satisfaction rates. In some cases, maintainer behaviours, planning deficiencies or outdated maintenance strategies contribute to the problem.
While these issues are not strictly Reliability Engineering problems, reliability teams can support effective Spare Parts Optimisation using statistical methods and criticality-based approaches. If minimum stock levels are set based largely on opinion, there’s a good chance the organisation is losing money.
Ageing Assets
Almost every asset-intensive organisation I work with is operating assets beyond their original retirement date. This is not necessarily a problem, but it does create risks that need to be managed.
The original retirement date should be a trigger to revisit assumptions, understand emerging risks and evaluate lifecycle costs before extending service life further. Reliability modelling and lifecycle costing are critical tools for developing this understanding and supporting informed business decisions.
Lack of Standardisation
Many organisations continue to support multiple equipment types that perform essentially the same function. The result is increased complexity in training, maintenance execution, and spare parts management. This is fundamentally a design issue, but it belongs here as many components of an asset – e.g. pumps, electric motors, and similar equipment items – can be standardised fairly easily across a plant or fleet of assets.
Reliability teams can help quantify the organisational costs associated with introducing and supporting additional equipment types and compare those costs against any performance benefits to determine if standardisation is a better option. Once again, lifecycle costing and reliability modeling play an important role in supporting these decisions.
No Continuous Improvement
This is probably my biggest bugbear. Too many organisations experience significant or recurring failures that are entirely preventable. Problems occur, operations are restored, and everybody moves on, but the root causes remain. Months or years later, the same failures return.
Defect Elimination is an essential capability for any asset-owning organisation. It provides a structured approach to identifying, analysing and addressing sources of failure. We routinely apply Root Cause Analysis to safety incidents because we recognise the value of preventing recurrence. We need to find the will to apply similar attention and discipline to reliability failures.
Start Where You Are
At the end of the day, “good” simply means being better tomorrow than you are today. Do that enough time and your organisation can become a leader in reliability excellence.
It doesn’t matter where you are in the lifecycle of your assets or how mature your reliability capability may be. There will always be opportunities to improve. Selecting the right ones requires an evidence-based assessment of current performance, a structured approach to diagnosing causes, and selection of interventions that address those causes.
The organisations that succeed are not necessarily the ones with the most sophisticated technology or the largest reliability teams. They are the organisations that resist the temptation to start with a solution and instead take the time to understand the problem first. From there, it is a question of disciplined execution of the improvement initiative, which is what I’ll address in the next article.
This article is part of the Reliability Engineering Excellence series by Assetivity. To read the other articles in this series, visit:
Article 1: Reliability Across the Lifecycle
