The Downtime Question That Can Change Your Automation ROI

Downtime Has to Be Part of the Automation Decision

The real question about automation downtime is not whether a robotic cell will ever stop. The question is how much production the plant can lose when it does. Operations also need to know how quickly the team can identify the cause and restore production.

Automation can concentrate production dependency. A manual process may allow other workstations to continue when one operator or station becomes unavailable. A robotic cell at a production constraint can create a different problem. One fault may interrupt several connected processes.

That does not make automation less reliable by default. It means the project team needs to treat downtime as a design and investment variable. Maintenance should not have to solve every recovery issue after commissioning.

A project may look attractive based on cycle time and labor savings. However, realistic interruption and recovery conditions can change that business case.


Start With the Production Consequence

There is no useful universal percentage for acceptable downtime. The same interruption can create very different consequences depending on the cell’s position in the production flow.

A downstream buffer may allow production to continue during a temporary stop. A robot feeding the plant’s main constraint may offer almost no equivalent protection. Every lost operating period can then reduce output that the plant cannot easily recover.

Start by connecting the cell to the production process rather than evaluating the robot alone. Ask what happens elsewhere in the plant when that cell stops.

Identify What Stops Producing

Determine whether a failure interrupts one operation, one machine, an entire line, or several downstream processes. Then identify the available alternatives.

Can operators perform the task manually? Can the plant redirect production? Can a buffer protect downstream operations? Can another shift recover the lost output?

Those answers change the economic impact of downtime. A plant with a practical bypass faces a different risk from one with a single point of failure.

Separate Lost Time From Lost Output

Machine downtime and production loss are related, but they are not always identical. A temporary stop may consume available buffer without reducing final output. Conversely, a short interruption at a tightly balanced bottleneck can create losses that continue after the original fault has been cleared.

Plants should therefore measure how a stop affects production. Track lost output, delayed orders, overtime recovery, idle equipment, scrap, rework, and additional interventions. These factors give management a more defensible basis for setting downtime limits.


Not All Downtime Creates the Same Production Risk

A single downtime figure can hide the information teams need to improve a robotic cell. Planned maintenance, recurring process stops, equipment failures, operator recovery delays, and upstream starvation all require different responses.

Before setting an acceptable downtime target, divide interruptions by cause. URT’s guidance on strategies to minimise downtime in robotic automation provides additional context for managing reliability after a cell enters production.

Planned Downtime

Maintenance teams can schedule inspections, controlled software work, tooling changes, and other planned interventions around production requirements. These activities still consume production time, but the plant can prepare people, parts, and production buffers before the stop.

Trying to eliminate all planned downtime can create additional risk. Instead, the plant should control these interventions and account for their production impact in the operating plan.

Unplanned Equipment Downtime

An unexpected robot, controller, tooling, sensor, safety-system, conveyor, or peripheral failure is different because the duration is uncertain. The robot arm is only one potential source of interruption. Cell reliability depends on the complete system, including end-of-arm tooling, fixtures, controls, communications, safety equipment, and connected upstream and downstream machinery.

This is why purchasing decisions should not evaluate robot reliability without evaluating the architecture around it. A reliable robot inside a poorly supported cell can still produce unacceptable downtime.

Process-Induced Stops

Some interruptions attributed to “the robot” are actually production-process problems. Inconsistent part presentation, damaged components, fixture variation, sensor contamination, upstream timing problems, or unexpected product conditions can prevent a correctly functioning robot from completing its sequence.

Automation does not remove these dependencies. If operators constantly correct process variation today, the automation project must address that variation directly. The cell design must either remove the cause or provide a controlled way to handle it.


Recovery Time Can Matter More Than Failure Frequency

Plants often concentrate on how frequently equipment might fail. Frequency matters, but it does not describe the full production risk. The other variable is how long the cell remains unavailable after an interruption.

An occasional fault can still create serious production risk when technicians need specialist support, difficult-to-source parts, software access, or extensive diagnosis. A recurring minor stop may create less disruption if trained staff can identify the problem and restore production safely.

This makes recovery capability part of the automation specification. The project should define what happens after abnormal conditions, not only how the cell behaves during its normal automatic cycle.

Design Recovery Logic Before Production Starts

Emergency stops, interrupted sequences, missing parts, machine faults, operator interventions, and power interruptions can leave connected equipment in different states. Poor recovery logic can turn a simple interruption into a long troubleshooting exercise.

The control architecture should help trained personnel identify where the sequence stopped. It should also show which conditions they need to restore before production can resume safely. This is one reason why the quality of the systems integration strategy affects operational risk after commissioning.

Internal Capability Changes the Downtime Calculation

Two plants operating similar robotic cells can experience very different consequences from the same fault. A site with trained technicians, documented recovery procedures, appropriate access, backups, and critical spare parts may restore production without waiting for external assistance.

A site without that capability may depend on a specialist for relatively routine problems. Before commissioning, manufacturers should define which interventions operators can handle, which require maintenance personnel, and which require external robotics or integration support.

This makes training a reliability investment rather than simply a commissioning task. Plants evaluating their support model should also define what training staff need to operate and maintain industrial robots before the cell becomes production-critical.


Downtime Risk Belongs in the ROI Model

An automation business case based only on purchase cost, labor savings, and theoretical cycle time is incomplete. The expected operating model should also account for availability, maintenance requirements, recovery capability, spare parts, support dependencies, and the economic effect of interruptions.

This does not require inventing a generic downtime allowance.

Do not use a generic downtime allowance. Instead, build the business case around the plant’s actual production exposure. Then test several downtime and recovery scenarios against that exposure.

Establish the Current Baseline

Measure the current process before estimating the automated one. Record downtime by cause, manual interventions, actual throughput, scrap, rework, changeover losses, maintenance demand, and recovery time.

This baseline prevents the team from comparing automation with an unrealistic version of manual production. It also shows whether the proposed cell addresses the loss that actually limits output.

Model the Consequence of Different Stop Conditions

The business case should examine what happens when the automated process becomes unavailable. Does production stop immediately? How much buffer exists? Is there a manual fallback? Can another machine absorb the work? Can lost output be recovered on another shift without creating additional cost elsewhere?

The purpose is not to predict every future failure. It is to identify how sensitive the investment is to availability and recovery assumptions. If a relatively modest deterioration in those assumptions destroys the projected return, reliability and support deserve more attention before approval.

Measure the Cell After Commissioning

Downtime should remain visible after the project has been accepted. Record the cause of interruptions rather than combining every stop into one availability number.

Useful measures include unplanned downtime by cause, recovery time after common stops, manual intervention frequency, throughput at the real production constraint, maintenance demand, scrap and rework, and performance against the original business case. URT’s article on KPIs for robotic automation explains how production metrics can be used to evaluate whether an installation is delivering the expected result.


The Support Model Should Match the Cost of a Stop

The cost of a production stop should determine the level of technical support. A non-critical robot with alternative production capacity may need a different support strategy from a cell that controls the plant’s main production constraint.

Match the support strategy to the production exposure. Review internal skills, documentation, diagnostic access, software backups, component availability, external support, and access to compatible replacement parts.

Critical Spares Should Follow Failure Consequence

Keeping every possible replacement component on site is rarely practical. Keeping none can also be expensive when a difficult-to-source component can stop a production-critical cell.

The spare-parts strategy should consider how essential the component is, whether another component can substitute for it, how quickly it can realistically be obtained, and how much production remains exposed while the cell is unavailable. This is especially relevant when controllers or peripheral equipment have long service histories or limited local support.

Documentation Is Part of Recovery Capability

Poor documentation makes a robotic cell harder to recover. The maintenance team needs access to electrical documentation, programs, parameter backups, tooling information, fault descriptions, and clear system responsibilities.

Do not allow recovery knowledge to depend on one technician or integrator. During project handover, give the people who will own the cell the information they need to troubleshoot and restore production safely.


When the Downtime Risk Is Too High to Approve the Project Yet

Sometimes a plant should delay automation even when a robot can technically perform the task. The process may suit robotics while the organization still lacks the resources to support the required availability.

Risk increases when the cell becomes a single point of dependency. An unstable process, unclear recovery procedures, limited maintenance capability, poor parts support, or no production fallback can increase that exposure further.

These conditions do not necessarily justify abandoning automation. Instead, redesign the project around the risk. Additional buffering, a bypass strategy, simpler tooling, better process control, clearer recovery logic, staff training, or improved spare-parts planning may reduce downtime exposure.

The same logic applies when automation is being justified mainly through aggressive utilization assumptions. A project that only works financially when the cell operates close to its theoretical production capacity leaves little room for maintenance, changeovers, process variation, or recovery. The business case should survive realistic operating conditions rather than depend on perfect ones.


What to Check Before Accepting the Downtime Risk

Use this checklist before project approval and again before final acceptance. Confirm that the plant understands what can interrupt production and how the team will respond when a stop occurs.

  • Define the production dependency: Identify which machines, processes, and customer output are affected when the cell stops.
  • Identify the true constraint: Determine whether the automated cell will sit at the operation that limits total production.
  • Measure the existing baseline: Record current downtime, recovery time, interventions, throughput, scrap, rework, and maintenance demand.
  • Separate downtime by cause: Distinguish planned maintenance, equipment faults, process-induced stops, upstream shortages, downstream blockage, and operator intervention.
  • Evaluate buffers and fallback: Determine whether production can continue temporarily, move to another route, or operate manually when the cell is unavailable.
  • Define recovery responsibility: Specify what operators, maintenance personnel, integrators, and external specialists are expected to handle.
  • Test abnormal conditions: Commission recovery from realistic interruptions rather than validating only the normal automatic cycle.
  • Review spare-parts exposure: Identify components whose unavailability could extend a production stop significantly.
  • Verify documentation and backups: Make sure the plant can access the information required to diagnose and restore the system.
  • Test the ROI assumptions: Check whether the investment remains defensible when realistic maintenance and downtime conditions are included.

If several of these questions cannot be answered before approval, the plant does not yet have a reliable downtime assumption. That uncertainty should be resolved before the expected productivity of the cell is treated as a financial result.


FAQ

How much downtime is acceptable for an industrial robot?

There is no universal acceptable amount. The answer depends on where the robot sits in the production flow, the output lost during a stop, available buffers or fallback capacity, and how quickly the plant can recover. Downtime should be evaluated through its production consequence rather than against a generic percentage.

Should planned maintenance count as downtime?

Yes. Planned maintenance still consumes potential production time, so include it when calculating realistic capacity. Track it separately from unplanned downtime because maintenance teams can schedule and prepare for these interventions.

Can a reliable robot still cause a high-downtime automation project?

Yes. The robot is only one component of the cell. Tooling, fixtures, sensors, safety systems, PLC logic, communications, conveyors, upstream equipment, downstream equipment, process variation, and operator recovery can all affect overall availability.

How should downtime be included in automation ROI?

Start with the current production baseline and then evaluate how interruptions in the proposed automated process affect actual output and operating cost. The model should include realistic maintenance, recovery, support, and production-dependency assumptions rather than assuming theoretical availability.

Does keeping more spare parts eliminate downtime risk?

No. Spare parts can reduce exposure to some hardware failures, but they do not solve process instability, programming problems, unclear recovery logic, inadequate training, or integration faults. Spare-parts planning should be one part of a wider recovery strategy.

When should an automation project be delayed because of downtime risk?

Consider delaying or redesigning the project when a production-critical cell lacks clear recovery procedures, support responsibilities, spare-parts planning, process stability, or a realistic fallback strategy. Address these conditions before relying on the cell’s projected output.


Talk to URT About Automation Downtime Risk

If you are evaluating automation downtime risk, contact URT. We will give you a direct, technical answer based on your actual production requirements.