From Imperfect Data to Lasting Results: How Reliability Improvements Gain Traction | Assetivity

In my previous article, I discussed how organisations can identify reliability improvement opportunities, understand the causes of poor performance, and select interventions that fit the problem. That’s an important first step, but identifying the right improvement is only half the battle. Achieving the improvement requires that the initiative gets implemented and this usually means overcoming a series of challenges, often from well-intentioned stakeholders. In this article, I’m going to address some of the most common blockers to change, starting with data.

Don’t Wait for Perfect Data

“We don’t have the data to make that decision.”

I regularly hear that statement and, if I’m being honest, I’ve occasionally said it myself. The problem is that it’s generally not true. More data, particularly good-quality data, will usually support a better decision, but that doesn’t mean we know nothing today, and it certainly doesn’t mean we should wait to act. Doing nothing is a decision in its own right. It locks in current performance and accepts today’s reliability outcomes as tomorrow’s outcomes. If we’ve already established that the current situation isn’t good enough, that’s rarely a sensible choice.

There’s also a second common fallacy in this space – the belief that uncertainty will disappear if we wait. The truth is that it won’t. Reliability engineering, like most business disciplines, operates in an environment of incomplete information. Assets change, operating conditions evolve, failure modes develop, and organisations learn new things over time. On top of this, we have Resnikoff’s Conundrum1, which reminds us that successful prevention of major failures often means there are very few examples available for analysis. The absence of historical failures doesn’t necessarily mean the risk is low. Sometimes it simply means previous efforts have done a good job of preventing them.

If waiting is not the answer, then we need to think about reliability improvement as a continuous improvement challenge. We make the best decision we can based on the information available today. Then we monitor the results, learn from what happens, and update the decision when better information becomes available. It’s simply Plan-Do-Check-Act applied to reliability decision-making.

I once heard a story that Apollo 11 was directly on course to the moon for less than two per cent of its journey. I’ve never found the original reference for that, but whether the exact number is right or not is beside the point. The mission succeeded because they continuously monitored and corrected the trajectory. If NASA had waited for perfection, they’d still be waiting.

To move forward with imperfect data, the first thing to recognise is that not all business decisions require the same level of certainty. We’re all familiar with the need to eliminate safety and environmental risks or to reduce them “So Far As Is Reasonably Practicable”. To meet that standard, we need to make decisions that are highly conservative – usually the left-hand tail of the distribution of possible values. This will be very sensitive to the quality of the data, as illustrated below:

We can think of this (without pretending to offer legal advice) as requiring our data to prove “beyond reasonable doubt” that our actions are safe and effective – but it also shows that we can make a decision even in the presence of poor-quality data. We just have to make sure it is sufficiently conservative – which is relatively easy when robust data collection and analysis techniques are used.

The situation improves further for other business problems, particularly cost optimisation problems, where a more reasonable standard of proof is “on the balance of probabilities”. In that case, both the high uncertainty and low uncertainty curves above support the same decision! In a real-world scenario, the decision would likely then be driven by other factors, such as the impact of failures on stakeholder reputation – or even the cost of collecting the required data.

The second thing to consider when dealing with imperfect data is that not all data needs to be purely statistical or drawn from your asset information systems. It’s often possible – and desirable – to make business decisions based on expert judgement. Your Reliability Engineers can apply Delphi and other elicitation techniques to rapidly develop defensible values for many parameters that you don’t have good history for and/or where it would be too expensive to gather that data. Humans can be surprisingly good at making these kinds of estimates when supported by the right techniques and one that I favour is the use of a “triangle distribution” to approximate a continuous curve by estimation of the maximum, minimum, and most likely values. The relationship between this and the unknown “true” distribution is shown below and can be used – with appropriate cautions – in the same way as other distributions.

Even if the quality of data forces a decision that is highly conservative – or prevents any decision at all – the attempt to make it can draw attention to the importance of the data. This can trigger an improvement action of its own – and even subtly influence the culture of record keeping, providing better data for the next attempt. Try getting the most vocal complainers together to make a decision and present them with their own data – full of blanks, default entries, and errors. They’ll quickly understand why data quality matters…

Obtain the Right Resources

The next challenge to implementation I’d like to address is resourcing. Many improvement programs assume the existing workforce can continue doing their normal job while simultaneously designing, implementing and embedding a new way of working.

Occasionally, that happens.

Usually, it doesn’t.

Meaningful change often requires additional effort and investment. For a period of time, organisations may need to support both the old and new processes simultaneously. Training needs to occur. Procedures need updates. Systems need modification.

That investment frequently creates a classic “J-curve” effect. Performance/expenditure may initially appear worse before the benefits emerge. Organisations that fail to anticipate this often become impatient and abandon worthwhile initiatives before they have been given a fair opportunity to succeed.

In many cases, the problem is us – the Reliability Engineering team. We can be so enthusiastic to make changes that we over-promise in order to get approval and then struggle to deliver within the available resources. Failure of the initiative is very likely, with the same kinds of risk to the reputation of Reliability Engineering that I discussed in the last article. Avoiding this risk requires a willingness to prioritise ruthlessly, as well as to communicate honestly about what can be achieved. 

There’s always more to be done than can be done, so less is more when it comes to selecting improvement initiatives. If you’ve spent the time to establish the link between Reliability Engineering and the organisation’s objectives, then you should be in a pretty good position to prioritise improvement initiatives. You might be able to manage two or even three initiatives simultaneously, but any more than that is almost certainly too much and will exhaust the available resources without delivering the required results. Show the rest on a roadmap so that people can see that there’s a plan but focus on today’s task. 

The honest communication is about ensuring that your stakeholders have a genuine understanding of the improvement initiative, potential benefits, and the time and other resources required to deliver them. It’s important to be realistic – don’t claim quick wins or underestimate the improvement effort. It’s easier to revisit a knock-back based on the truth than to justify a failed initiative based on over-promising.

Building Momentum Through Visible Progress

Once you’ve got your resources, you can help yourself – and your stakeholders – by making progress visible. This is a significant challenge for Reliability Engineering because the stochastic nature of failure means that RAM benefits invariably take longer to realise than other improvements. They can even be invisible where the initiative is preventing failures from occurring. If the organisation is not prepared for this, it can lead to stakeholder disillusionment and, more often than it should, abandonment of a sensible initiative. 

Since most improvement initiatives should run as formal projects, there’s obviously a lot of good practice project management that is important to success in this space – and there are others better placed than me to advise on many of them. From a Reliability Engineering perspective, my key concerns are related to the measurement of success, use of pilot studies, and the need for on-going stakeholder engagement.

Firstly, the challenge in seeing RAM improvements needs to be addressed by setting the right measures of success. Outcomes (cost reduction, availability improvement, etc) need to be measured, but these will take considerable time – probably well beyond the scope of a typical business improvement initiative. Accordingly, they need to be included in the ongoing business performance measures, not the success measures for the initiative. These should focus on implementation of progress and leading RAM measures such as process adoption and implementation of changes.

Secondly, I think that many organisations are misusing pilot studies. These are a great way to start an improvement initiative, but they can also be a risk to success. A “pilot study” implies uncertainty – it is often a test of the capability (people, process, systems, data) and it’s therefore reasonable to expect that it will take longer, deliver less, and quite possibly need reworking. Even if the uncertainty to be tested was the size of the benefit, the setup costs for a pilot are probably a substantial fraction of those required for a full roll-out. Unfortunately, many pilot studies are treated as an opportunity for “quick wins”, which leads to disappointment and, often, a “no-go” decision on what is a valuable initiative. These factors need to be considered in setting the success criteria for a pilot.

Finally, the nature of RAM improvements means that the project plan needs to be established for a marathon, not a sprint. It’s a long-term commitment that requires ongoing stakeholder engagement and reinforcement. Human beings aren’t really good at that kind of commitment, so it’s important to make sure that this is part of the plan.

Progress Beats Perfection

At the end of the day, reliability excellence is never the result of a single initiative. It’s built on making better decisions, learning from outcomes, and continuously adjusting courses. If you don’t think your organisation is ready for a long-term commitment, then it might be better to take a Defect Elimination approach, which will be the topic of a future article. Before that, however, I want to explore the details of how to build a Reliability Engineering capability.

To read the other articles in this series, visit:

Article 1: Reliability Engineering Excellence: Reliability Across the Lifecycle

Article 2: Where Should Your Reliability Improvement Journey Begin?

Back to top