$50B a year in unplanned downtime is a capacity problem

Manufacturing loses roughly $50 billion a year to unplanned downtime, and about 15 hours a week per facility sits idle. Almost every attempt to fix this has been an attempt to see the problem sooner, more sensors, better condition monitoring, predictive maintenance models. The industry has spent a decade getting very good at detection.
And the downtime is still there. That should tell us the bottleneck was never visibility.
What happens between the alert and the fix
Follow a single anomaly through a plant. A vibration signature drifts. An alert fires. Now someone has to decide whether it matters, check whether this line is scheduled for a changeover anyway, look up whether the spare is in stores or on a six-week lead time, work out which shift can take the machine down without breaching a customer commitment, raise the work order, and chase the vendor.
None of that is detection. All of it is coordination across an ERP, a maintenance system, a production schedule and three inboxes — and all of it lands on a small number of experienced people who are already fully committed. The alert was free. The response is the scarce resource.
Why predictive maintenance underdelivers
A predictive model tells you a bearing has perhaps three weeks left. Useful — but only if someone acts inside those three weeks. If the plant is already running at capacity on coordination, the prediction joins a queue, and the machine fails on schedule anyway. The model was right and the outcome was unchanged.
This is the same pattern that leaves 95% of GenAI pilots with zero P&L impact. Systems that advise get deployed; the action still falls on the team. Where there is no spare human capacity, better advice produces no better result.
The plant knowledge that is not in the plant systems
There is a further complication, and it is the one that defeats scripted automation. The right response to an anomaly depends on knowledge that is nowhere in the MES or the ERP: that this asset always reads high after a wash-down, that this supplier quotes six weeks and delivers in three, that this customer will accept a two-day slip but must be told today.
That knowledge sits with the maintenance planner who has been there nineteen years. It is why rule-based workflow tools break on contact with a real plant: they encode how the work is done without the reasoning that decides when the rule should not apply. And it is why the retirement curve matters here more than in most industries — roughly 10,000 experienced professionals retire every day, and plants are disproportionately exposed.
Treat it as a capacity problem
Reframed as capacity, the question changes usefully. Not “how do we predict failures better” but “who is going to do the coordination work when an anomaly appears at 2am on a Sunday.” An autonomous operator answers that literally: it reads the anomaly, checks the schedule and the stores, weighs the customer commitment, raises the work order, contacts the vendor, and escalates to a person only when the judgement genuinely requires one.
That is a different proposition from a dashboard, and it is measured differently. Not model accuracy — downtime hours avoided, work orders raised without a human, and hours returned to the planners who were the constraint in the first place.
What this looks like in practice
It runs on the systems you already have. Evos connects to the MES, ERP, maintenance system and email you run today, and puts an operator on one workflow — live in under 24 hours, with no migration and no rip-and-replace. It earns autonomy the way a new hire does: recommending first, acting once it has proven it decides the way your planners would.
The $50 billion is not lost to machines breaking. Machines have always broken. It is lost in the hours between the alert and the fix, where there is nobody free to act. Book an assessment and we will show you where those hours are going in your plant.
