Skip to main content
Transformidy

Article

Automation Is Only As Good As Its Failure State

The CBSA airport kiosk outage shows why organizations must design fallback capacity, communication and recovery as part of automation itself.

Published
September 13, 2026
Updated
September 13, 2026
Reading time
7 min
Matte paper aviation recovery illustration showing an automated journey breaking into manual processing, queues and downstream recovery paths.

Article Body

Automation is usually designed for the state everyone hopes will happen.

Matte paper system map showing failure detection, fallback capacity, customer communication, recovery and learning.
What changes

The customer arrives. The kiosk works. The identity check completes. The queue moves. The transaction clears. The system produces speed, consistency and scale.

The more important test is what happens when that normal state disappears.

On Friday, September 11, a widespread outage of Canada Border Services Agency primary inspection kiosks disrupted arrivals at Canadian airports, including Toronto Pearson and Montréal-Trudeau, according to the Sunday intelligence packet. The outage was reported in the afternoon and restored later that evening. Travellers were moved to manual processing, with reported lineups and waits extending beyond ordinary expectations. The same evidence packet also points to a separate Pearson kiosk outage in September 2025 that contributed to arrival journeys approaching four hours.

That is not only an airport story.

It is an automation story.

The Revenue Unknown is:

How much customer, operational and reputational value is exposed because organizations design automated capacity without adequately designing fallback capacity?

The successful state of automation is easy to diagram:

traveller -> kiosk -> verification -> clearance

The real operating model also needs the failed state:

kiosk unavailable -> Recognition -> capacity shift -> manual process -> traveller communication -> recovery -> downstream consequence

If the second path is weak, the first path is more fragile than it appears.

The Fallback Is Part Of The Product

Many organizations treat fallback as a contingency plan, not as part of the experience.

That is the mistake.

For the customer, the experience does not divide neatly into normal operations and incident response. The experience is the whole path. If the automated state is fast but the failure state is confusing, slow and under-resourced, then the organization has not built a resilient digital experience. It has built a fast experience with a hidden cliff.

A fallback that technically exists but cannot absorb actual demand is not meaningful resilience. Manual processing may be available, but if it cannot handle the volume created when automation stops, the customer still experiences a breakdown. A support phone number may exist, but if the queue explodes when the app fails, the recovery path is not real. A branch may be open, but if customers have already been trained into digital-only journeys, the human channel may not have the staffing or authority to recover the event.

This applies far beyond border kiosks.

Retail self-checkouts need a designed failure state. Banking apps need one. Airline apps need one. Payment systems need one. AI agents need one. Government-service portals need one. Real-time settlement needs one. Every system that removes friction in the normal state has to answer what happens when the normal state cannot continue.

The stronger Transformidy principle is:

Automation is only as resilient as the experience it provides when automation stops working.

Efficiency Can Hide Capacity Risk

Automation often creates a convincing efficiency story.

It reduces labour per transaction. It increases throughput. It standardizes decisions. It lets customers self-serve. It allows an organization to handle more demand without adding equivalent human capacity.

Those gains are real.

But the efficiency case can hide a capacity decision.

If the organization reduces manual capacity because automation now carries the normal workload, then the fallback capacity may become symbolic. It is present on the org chart or in the incident manual, but not large enough to carry the real demand when needed.

That creates a second Revenue Unknown:

What level of fallback capacity should an organization preserve for systems expected to work almost all the time?

The answer cannot be zero. But it also cannot be infinite. The practical question is how to decide.

Leaders need to evaluate the cost of preserving dormant or partially used recovery capacity against the value at risk when automated systems fail. That means looking beyond uptime. Uptime tells leaders how often the system is available. It does not show whether the organization can protect customers when the system is unavailable.

The better measures include degraded-state throughput, queue duration, staff reassignment time, customer communication speed, exception handling, recovery completion, downstream missed commitments and time to full service restoration.

Failure-State Architecture

Transformidy should treat this as an Experience Failure-State Architecture problem.

The method is simple:

automated state -> failure detection -> decision deadline -> fallback capacity -> communication -> recovery -> downstream consequence -> learning

Each step exposes a design decision.

Failure detection asks how quickly the organization knows the system is down and who has authority to act. Decision deadline asks how long leaders have before the customer impact becomes materially worse. Fallback capacity asks what resources can absorb demand. Communication asks whether customers know what is happening, what to do and what not to do. Recovery asks whether the organization can move customers through the failed state without creating further dead ends. Downstream consequence asks what the outage triggers elsewhere: missed connections, delayed baggage, hotel arrival problems, refund claims, staffing pressure, partner confusion or trust loss.

Learning is the stage many organizations miss.

After the system returns, the organization often closes the incident. The more useful question is whether the outage changed future design: staffing models, manual processing capacity, customer messaging, dependency mapping, technology redundancy, escalation authority or partner coordination.

If not, the same breakdown can repeat under a different label.

The Same Problem Is Coming For AI Agents

The kiosk outage is also a useful preview of where agentic systems are heading.

As organizations allow AI systems to recommend, decide, transact, route, summarize, approve or recover more work, they will be tempted to measure the successful path. Did the agent answer? Did it complete the form? Did it resolve the contact? Did it route the payment? Did it book the trip? Did it finish the workflow?

Those questions are necessary, but incomplete.

The harder questions are about the failed path. What happens when the agent is uncertain? What happens when the source data is unavailable? What happens when the customer disputes the action? What happens when the agent reaches the edge of its authority? What happens when the handoff queue is full? What happens when the customer needs a human but the human channel was reduced because automation was expected to carry the work?

That is the same failure-state problem in a different form.

Automation does not eliminate the need for human capacity, judgment or recovery. It changes where that capacity is needed and how quickly it must appear. The more autonomous the system becomes, the more explicit the fallback architecture must be.

This is why failure-state design belongs in executive governance, not only in incident response. Leaders should know which automated journeys have degraded-state playbooks, what those playbooks can actually absorb, and which customer promises become exposed when the system fails.

The Playbook Is A Design Object

The practical response is not to reject automation.

It is to design the failed state with the same seriousness as the success state.

An effective playbook should define the trigger, the owner, the decision deadline, the fallback capacity, the customer communication, the partner notification, the exception authority and the learning requirement. Each element should be tested under realistic load, not only documented.

For border kiosks, that means understanding how many travellers can be manually processed per hour under different staffing conditions. For banking, it means knowing how customers regain access when digital identity fails. For retail, it means knowing how stores recover when self-checkout, payments or inventory systems go down. For airlines, it means knowing how an app outage changes check-in, standby, rebooking and customer-care demand.

The failed state should not be the moment when the organization starts inventing the experience.

Experience The Skies Spotlight

Airports make failure-state design especially visible because many organizations share the journey.

When border-processing automation fails, the passenger does not experience only a CBSA outage. They experience an arrival delay, a queue, uncertainty, possible missed connections, baggage timing complications, ground-transport changes, hotel check-in pressure, airline-service ambiguity and family or business disruption.

That makes the aviation Revenue Unknown sharper:

How much fallback capacity should airports preserve for systems that are expected to work almost all the time?

Experience The Skies should examine border processing, passenger flow, aircraft holds, missed connections, baggage complications, hotel arrival impacts and downstream vendor effects as one recovery system. The goal is not to assign blame to a single participant. The goal is to recognize that a fragmented ownership model can make the passenger recovery path weaker than any one organization's technical response.

The stronger airport metric is not only kiosk uptime. It is passenger throughput and recovery quality under degraded operating conditions.

Sources

  • Transformidy Sunday Intelligence packet, September 13, 2026, CBSA airport kiosk outage object. Source verification required before publication.
  • CBSA public communications, pending direct source capture.
  • PAX airport disruption reporting, pending direct source capture.

FAQ

What is the main idea of Automation Is Only As Good As Its Failure State?

Automation Is Only As Good As Its Failure State explains a change leaders should not treat as background noise. It shows what evidence is visible, what may be changing underneath it, and which decision window remains open.

Why does Automation Is Only As Good As Its Failure State matter for Experience Intelligence?

The article helps readers see how an experience, relationship, capability, or value condition may be changing before the consequence is fully visible.

What Revenue Unknown does this article help identify?

It frames the unresolved commercial or operating question created by the change: what value, risk, hidden demand, relationship movement, or capability gap may exist but has not yet been measured or decided.

How should leaders use this article in the Special Intelligence series?

Use it as a prompt to separate observed evidence from interpretation, name the decision that still has to be made, and identify what would validate whether the interpretation is right.