1. The Situation
Years ago, I wrote about the biggest mistakes companies make with reliability engineers and why their numbers seem to increase but the equipment reliability does not continually improve.
Over the years we have built reliability teams. They are smart, capable, and well trained. They do RCAs and can hardly keep up with the demand so they create RCA trigger rules to limit demand and focus on high value projects. Improvements are implemented and monitored to ensure they remove those breakdowns. But years later… not much has changed, breakdowns are still happening and the same or similar chronic issues keep coming back and the overall uptime of the equipment has not improved and in some cases gotten worse.
And quietly, people start thinking: “The reliability team isn’t delivering.”
But when you look closer, the problem isn’t the capability of the reliability team. It’s the ownership. The execution team keeps doing maintenance and the reliability team keeps doing analysis and somewhere in between… the improvements never significantly move the needle.
2. What Most Companies Do
Most organisations structure it like this:
- Reliability team = responsible for improvement
- Execution team = responsible for maintenance and fixing breakdowns
So when a failure occurs:
- The execution team fixes it and moves on
- The reliability team analyses it later (or tries to without adequate information)
This creates a gap. The people closest to the failure don’t own the improvement and the people doing the analysis don’t own the execution.
So we end up with:
- Good reports
- Logical recommendations
- Limited sustainable improvement
3. What Actually Works
The accountability for reliability improvement must sit with the maintenance execution superintendent.
They own:
- The fleet or plant area
- The supervisors and execution teams
- The work execution standards
- And most importantly, the outcomes
I developed the breakdown response model some years ago to help teams understand how to get practical outcomes and work together.
For each breakdown (or at least the top 5 to start), the execution team must ask:
- Should this have been detected earlier?
- Was there already a task to prevent it?
- Was it executed properly?
- Was this a known defect that wasn’t prioritised and actioned?
Then act immediately to close the gap.
Only when the cause isn’t clear, or everything was done correctly and it still failed (only ~25% of cases in my experience), should the superintendent engage the reliability team.
The reliability team is there to support, not own.
They should bring:
- Structured analysis processes
- Failure mode thinking
- Technical depth
But they should be used when needed, not as the default response.
4. Clear Principle
You cannot outsource reliability improvement, Reliability is built through execution and execution sits with the superintendent and their team.
Execution teams don’t need to be experts in every tool. They just need to:
- Understand the basics
- Follow a simple response process
- Know when to bring in support
That’s enough to eliminate the majority of failures.
In your operation — are you solving breakdowns where they happen… or handing them off to be analysed later?
In my maintenance superintendent coaching program, we focus on exactly this, helping superintendents take ownership of reliability and consistently achieve their availability and performance targets.






