Ask most plant managers for their reliability numbers and they'll hand you MTBF and MTTR broken out by line. Line 3 versus Line 5, section by section, month over month. That's not wrong. It's just not where the money is.
Mean time between failures and mean time to repair are two of the oldest metrics in industrial maintenance, and container glass plants have been tracking them for decades. The problem isn't the metric. It's the axis you're slicing it on. A line-level MTBF of 340 hours tells you the line is unreliable. It doesn't tell you whether the culprit is plunger drift, a cracked blank mould, a choked feeder tube, or a pusher bar that's been slightly out of time since the last mould change. Average those causes together and you get a number that looks stable month to month while the actual failure mix underneath is shifting entirely.
Line-level averages hide where the losses actually live
I worked a two-furnace, six-line plant in the GCC in 2017 where the maintenance report showed MTBF holding steady around 280 hours for 18 straight months. Steady looked like control. It wasn't. When we broke the same downtime events down by failure mode instead of by line, one cause, recurring plunger seal failures on two specific sections, accounted for 41% of unplanned stops across the whole plant. Nobody had seen it because it was buried inside six separate line averages, each one diluted by everything else that happens on a forming line in a month.
That's the pattern on nearly every audit I run. Line-level MTBF and MTTR are the numbers that get reported upward because they're easy to pull from the DCS or the CMMS. Failure-mode MTBF is the number that actually changes behaviour, because it points at a specific cause with a specific owner and a specific fix. Until you code downtime by mode, you're managing a symptom.
What failure-mode coding looks like on a real shift
This isn't complicated, but it does need discipline the CMMS won't enforce for you. Every stoppage gets a cause code at the point of entry, not reconstructed from memory two days later. Stuck plunger. Baffle misalignment. Swab burn on the blank. Choked orifice ring. Take-out arm timing drift. Feeder tube crack. Each mode gets its own MTBF and MTTR line, tracked section by section, not lumped into "mechanical" or "process" as a catch-all.
The hot-end superintendent owns the cause code at the point of entry, not the operator, who's busy running the section, and not maintenance planning, who wasn't standing at the machine when it happened. Get that ownership wrong and the codes drift toward whatever's fastest to type, and "mechanical, other" becomes 60% of the downtime log within a quarter. I've seen that exact drift on a US plant running older Emhart IS machines with relay-era control panels and no section-level fault logging beyond a red light (and yes, I know the shift supervisor will swear the panel logs everything, it doesn't). Every stop got coded "unknown" because nobody had built the habit of writing down what actually happened before the section was running again.
On mould-related failures specifically, tie the cause code back to the mould's own maintenance history: preheat curve conformance, cycle count since last shop visit, cast iron versus bronze plunger tips. A blank mould that cracks at 60,000 cycles isn't the same failure mode as one that cracks at 12,000, especially if the preheat curve wasn't holding to target, 480°C ±10°C is the number I look for, before it went back into rotation.
Persistent failure modes need an action plan, not another work order
Here's where most plants actually fall down. They get the coding right, they build the Pareto chart, they identify that plunger drift or delivery system chokes are the repeat offender, and then they issue another work order and move on. A work order fixes an instance. It doesn't fix a pattern.
A persistent failure mode, one that shows up three or more times in a rolling 90 days on the same section or the same mould set, needs a named owner, a root-cause investigation, and a closed-loop check that the fix actually held. That means:
- A documented root cause, not a guess written up after the fact
- A single accountable owner with a date, not "maintenance to review"
- A defined re-check interval to confirm the MTBF for that mode actually moved
- A record kept against the SKU or mould set, not just the machine
And if the MTBF for that mode doesn't move after the fix, the action plan gets escalated, not closed. That's the discipline most plants skip. They treat the action plan as a documentation exercise instead of a loop that has to close.
A failure mode you've coded but never closed the loop on isn't data. It's a diary.
The handover gap that erases the data before anyone sees it
The 0600 handover is where failure-mode data usually dies. I see it on most audits: night shift logs the stop, maybe scribbles a cause on a whiteboard or a paper sheet, and by the time day shift's supervisor reviews overnight events, the detail that would have let you code it correctly is gone. Was it a cold mould after the changeover, or a genuine crack? Was the plunger drift new, or had the operator already compensated for it twice that shift without logging either instance? Nobody asks, because the handover conversation is built around what's running and what's not, not what failed and why.
Look, the coding only works if someone actually believes it's worth doing right. Fix the handover template before you fix the CMMS. Five minutes of structured cause capture at shift change is worth more than a year of "mechanical, other" entries you'll never unpick.
This is a management discipline, not a maintenance department problem
Reliability data only changes anything if the people reviewing it have the authority to act on it, and that's a management structure question before it's a technical one. Not a maintenance metric. A management one. Our management audit looks specifically at whether the meeting cadence and reporting lines let failure-mode data actually reach someone who can approve the fix. A brilliant Pareto chart that stops at a maintenance planner's desk changes nothing.
Zaid Hassoneh built this discipline the hard way, tracking failure modes on the floor at O-I Brisbane from 2005 and carrying it through to the Arglass Yamamura greenfield build. It's the same logic behind the KPI tracking inside our Job Change Tool: track by named cause, not by shift and not by line average, and the pattern shows up in weeks instead of years. As a vendor-neutral container glass consultant, we don't care which OEM's control system logged the stop. We care whether the cause code is honest and whether someone owns the fix.
If your reliability reporting still lives at the line level, start smaller than you think. Pick your worst-performing section, code every stop by failure mode for 30 days, and see what's actually hiding under the average. Our hot end audit does exactly that as a starting baseline, before we ever touch the mould shop or the batch house.
That GCC plant with the steady 280-hour MTBF never touched its plunger seal problem until the number was broken open by failure mode. Once it was, MTTR on that specific mode came down -58% inside two quarters. The line average barely moved. The plant did.