Skip to main content
Transformidy

Article

One DNS Failure Exposed a Much Bigger Cloud Risk

An AWS DNS failure in October 2025 showed how a narrow infrastructure problem can create a much wider business blast radius. Forrester expects at least two more multiday cloud outages in 2026.

Published
July 22, 2024
Updated
June 18, 2026
Reading time
8 min
Editorial illustration for One DNS Failure Exposed a Much Bigger Cloud Risk

2026 updated analysis

What changed since the original article

This page keeps the original Transformidy article as the canonical record and leads with the current interpretation, source notes, and Revenue Unknown framing.

The Failure and the Disproportionate Blast Radius

AWS reported that the October 20, 2025 event involved DNS resolution issues for regional DynamoDB service endpoints in US-EAST-1. The initial DNS issue was mitigated early, but related subsystem impairment meant full normalization took longer. Separately, third-party outage reporting estimated that the downstream impact reached more than 3,500 companies across more than 60 countries and generated more than 17 million Downdetector reports.

The confirmed root issue was narrow relative to the observed business blast radius. This is the pattern worth sitting with rather than skipping past on the way to the bigger headline numbers. A failure does not need to be dramatic or sophisticated to cause damage at this scale; it only needs to sit at a sufficiently concentrated point in a widely shared dependency chain.

AWS was not alone in October and November 2025. Roughly a week later, Microsoft's Azure, Microsoft 365, Outlook, and Xbox Live faced a massive global outage beginning October 28, 2025, an incident traced to an Azure Front Door configuration change that generated more than 18,000 user reports at its peak. Two major hyperscalers, roughly a week apart, both taken down by relatively narrow, internal technical failures with disproportionately broad customer-facing consequences.

AWS outage, October 20, 2025
3,500+ / 60+

Estimated companies affected across countries, from a narrow shared-infrastructure failure

More than 15 hours to full restoration; over 17 million Downdetector reports generated.

Why Forrester Expects This to Keep Happening

Forrester's assessment of why this pattern is likely to continue is specific. Principal analyst Lee Sustar expects "at least two major multiday outages in 2026," connecting the risk to hyperscalers prioritizing AI-oriented infrastructure while older environments carry growing complexity.

That framing identifies a structural, not incidental, cause. Cloud providers are actively reallocating investment toward AI infrastructure, which means the legacy systems underlying most current enterprise workloads are, per Sustar's assessment, receiving comparatively less attention even as their operational complexity grows. Analyst Catalin Voicu's related observation reinforces the systemic read: organizations are recognizing "the fragility of centralized cloud infrastructure and how a single issue can cascade into multiple failures."

For a business, the practical implication is that this is not a one-time incident to review and move past. It is a forecast, from an independent analyst firm, of a continuing pattern with a specific, named structural driver. A resilience plan built to address October 2025's specific outage without addressing the broader "aging infrastructure under growing complexity" dynamic Sustar describes is solving last year's incident, not preparing for next year's, which Forrester's own analysis suggests is already likely.

What changes

Reviewing the Last Outage vs. Preparing for the Next One

One patches a specific known failure. The other assumes the underlying pattern will repeat.

Reviewing the last outage: documenting exactly what happened in October 2025 and confirming the specific DNS issue is resolved, useful but narrowly scoped. Preparing for the next one: building and testing failover capability against Forrester's forecast of at least two more multiday outages in 2026, treating the underlying dependency-concentration risk as ongoing rather than resolved.

"We believe this strategy will have some meaningful fallout in the form of at least two major multiday outages in 2026."

Lee Sustar, Principal Analyst, Forrester

The Leadership Move

The structural choice for any technology or business continuity leader is whether to treat the October 2025 outages as a closed, one-time incident, or as confirmation of an ongoing, forecasted risk pattern requiring a standing response.

Ownership

Infrastructure and business continuity leadership own the responsibility to map which customer-facing functions depend on a single cloud provider or region, and to build tested, not just documented, failover capability, given Forrester's specific forecast of continued multiday outages.

Tradeoff

Building genuine multi-region or multi-provider redundancy costs meaningfully more than relying on a single provider's infrastructure. The tradeoff against that cost is exposure to outages at the scale AWS's October 2025 incident demonstrated, roughly $9,000 per minute at the current median enterprise rate, sustained over many hours.

Human consequence

End customers of the more than 3,500 affected companies experienced downtime with no understanding of or connection to the actual cause, a DNS record at a cloud provider several layers removed from the service they were trying to use, and no ability to distinguish a well-prepared company's outage from a poorly prepared one in the moment.

Next Move

If your business runs primarily on a single cloud provider or region: Map which customer-facing functions have no tested failover, and treat closing that gap as a standing priority given Forrester's forecast of continued multiday outages in 2026.

If you already reviewed the October 2025 incident internally: Confirm the review addressed the structural pattern, aging infrastructure under growing complexity, not just the specific DNS failure, since the forecasted risk is broader than any single root cause.

FAQ

What caused the October 2025 AWS outage, and how big was its impact?

AWS traced the October 20, 2025 outage to DNS resolution issues for regional DynamoDB service endpoints in US-EAST-1. AWS reported that the initial DNS issue was mitigated early, while full service normalization took longer because related subsystems remained impaired. Third-party outage reporting estimated broad downstream impact across thousands of organizations and many countries.

Is this an isolated incident or part of a broader pattern?

Part of a broader pattern. Forrester principal analyst Lee Sustar expects "at least two major multiday outages in 2026," attributing the trend to hyperscalers diverting investment away from legacy infrastructure toward GPU-centric data centers for AI workloads while aging infrastructure faces growing complexity. Microsoft's Azure, Microsoft 365, Outlook, and Xbox Live also experienced a major global outage beginning October 28, 2025, roughly a week after the AWS incident.

Why did such a narrow technical failure cause such a large business impact?

Because many unrelated companies and services depend on the same underlying cloud region and infrastructure without necessarily understanding the full extent of their shared dependency. The technical trigger can be narrow while the customer-facing blast radius is broad, which makes concentrated cloud dependency a systemic risk rather than an isolated vendor problem.

What should a business actually do in response to this pattern?

Map which customer-facing functions depend on a single cloud provider or region without a tested failover, and treat that map as a live risk register rather than a one-time audit, since Forrester's own analysis suggests the underlying causes are structural and unlikely to resolve in 2026.

Sources & References

Original article archive

Original article published July 22, 2024: "Outages - Safeguarding Customer Experience ". Preserved here for provenance, historical context, and citation continuity.

In today's digital age, customer experience (CX) is paramount. Businesses rely on technology to deliver seamless and reliable experiences for their customers, but recent outages like the one experienced by Microsoft and CrowdStrike serve as stark reminders of the fragility of our interconnected systems. These outages can have a devastating impact on businesses, disrupting operations, damaging brand reputation, and ultimately leading to lost revenue.

https://transformidy.com/insight/microsoft-crowdstrike-outage-customers/
Were you affected the by Microsoft/CrowdStrike outage that impacted million of customers and businesses on Friday, July 20, 2024?

The Microsoft and CrowdStrike outage left millions of end users without access to critical services. Businesses across various industries from finance and healthcare, retail and manufacturing, air travel to hospitality were affected. This insight continues our previous discussion and highlights some proactive measures for companies to safeguard customer experience in the face of potential outages.

In-depth Understanding Of The Outage's Impact On Customer Experience

Outages can disrupt customer journeys in a number of ways. When applications go down, customers cannot access the products or services they rely on (e.g., access mobile application to check-in or receive flight information, purchase a ticket online for a concert, scheduling software for surgery and patient care, receive funds for a loan). Loss of access can will lead to frustration and anger on a smaller scale but could also need to business failures and physical harm. In the case of the Microsoft and CrowdStrike outage, businesses that relied on these services for core operations were brought to a standstill.

A computer screen with a blue screen on it
Is your company prepared for the fallouts in customer experiences through an outage? Photo by Milad Fakurian on Unsplash

Beyond immediate disruptions, outages can also damage brand reputation. When customers experience problems, they expect them to be resolved in a reasonable amount of time. They may be less likely to trust the business in the future if this expectation is not met. Customer churn and lost revenue is an outcome businesses do not want or can afford.

How To Safeguarding Your Company's Customer Experience Through Outages?

In light of these potential consequences, businesses must take a proactive approach to safeguarding customer experience. Here are some key strategies to consider:

  • Implement Robust Risk Management: A proactive risk management strategy is essential for identifying and mitigating potential threats to your applications and systems. This includes identify key customer experience applications and touch points, conducting regular risk assessments on how a reduction or removal of service would impact customers, employees, and other business stakeholders, implementing measures to reduce outage possibility, and crafting and testing a disaster recovery/business resumption plans. 

    By anticipating potential problems, companies can minimize their impact on your customers.
  • Improve Application Redundancy: Redundancy is the key to ensuring that key customer or employee applications remain available even in the event of an outage. This means having backup systems in place that can take over if your primary systems fail. The best approach for your business will depend on your specific needs and budget.

    AI can be a valuable tool. By analyzing vast amounts of data, AI can predict potential issues and enable preventive maintenance, minimizing downtime. However, it's crucial to train AI models on high-quality data and continuously monitor them to avoid biases or errors. Explainable AI can also enhance trust in its decision-making capabilities.
  • Prioritize Rigorous Testing: Regular testing is crucial for ensuring the reliability and performance of key customer and employee-facing applications. This includes testing for functionality, performance, and security. By identifying and fixing potential problems before they occur, you can help to prevent outages and safeguard your customer experience. Companies should also communicate any potential service outage resulted from key software updates or fixes being implemented. This extends to any third party involvement in carrying out these changes as customers do not always know difference and will uphold companies to act in the best regards.
  • Establish Effective Communication Channels: Communication is key during an outage. When an outage occurs, it is important to communicate with your customers, employees, and key stakeholders quickly, transparently and in the channels they prefer. Provide them with updates on the situation, what/how the company is doing to resolve the issue, and an estimated time for restoration. By keeping different parties informed, companies can help to minimize frustration and maintain trust. The marketing, legal, and public relations team should be on hand to manage any fallout from extended outages.
  • Invest in Customer Experience Monitoring: Proactive customer experience applications monitoring can help companies identify potential issues before they become outages. By tracking key metrics such as application performance and customer sentiment through social media/operational performance/feedback monitoring, companies can gain insights that can help prevent problems and improve overall customer satisfaction.
  • Embrace a Culture of Continuous Improvement: Safeguarding customer experience is an ongoing process. It is important to continually evaluate approved strategies and make improvements, as needed. Fostering a culture of continuous improvement with key stakeholders across different departments will ensure that the business is always prepared to deliver the best possible experience for customers, employees, and other stakeholders. Collaboration is key to success!
Group of people using laptop computer / Collaboration on customer-facing applications and software is key to managing outage impacts
Collaboration on customer-facing applications and software is key to managing outage impacts Photo by Annie Spratt on Unsplash

Building Resilience For The Future

By following these strategies, businesses can build resilience against outages and safeguard their customer experience. In today's digital world, where outages are a constant threat, taking a proactive approach is essential for ensuring business continuity and success. 

By prioritizing application redundancy, risk management, testing, communication, and customer experience monitoring, businesses can create a more reliable and customer-centric environment.

By understanding what are the paths forward to manage fallouts of outages, companies will potentially improve engagement/loyalty, and protect revenue.

Chief Experience Officer/Chief Customer Officer's Role In Outages

As the public face of the organization, their ability to communicate effectively, empathetically, and transparently is paramount. Beyond crisis communication, the CXO/CCO must orchestrate a comprehensive response, including providing immediate support to affected customers, mitigating reputational damage, and closely monitoring customer sentiment across various channels.

A critical aspect of their role is post-outage analysis. This involves a deep dive into the company's response to identify areas for improvement in crisis management, communication protocols, and operational resilience. By understanding the root causes of the outage and its impact on customers, the CXO/CCO can develop strategies to prevent similar incidents and enhance the company's ability to respond effectively in the future.

Moreover, rebuilding customer trust is essential. The CXO/CCO must spearhead efforts to regain customer confidence through tangible actions such as compensation, loyalty programs, or exclusive benefits. This demonstrates the company's commitment to customer satisfaction and reinforces its dedication to preventing future disruptions.

Transform For The Better

15-Point Checklist for Safeguarding Customer Experience

To effectively implement the strategies outlined above, consider the following checklist:

Application Redundancy

  1. Identify critical applications: Determine which applications are essential for business operations.
  2. Develop a redundancy strategy: Choose the appropriate redundancy level (hot, warm, cold) for each critical application.
  3. Implement backup systems: Ensure backup systems are in place and regularly tested.
  4. Test failover procedures: Conduct regular failover tests to verify system performance.

Risk Management

  1. Conduct regular risk assessments: Identify potential threats and vulnerabilities.
  2. Implement security measures: Protect systems and data with appropriate security controls.
  3. Develop a disaster recovery plan: Outline steps to be taken in case of an outage.
  4. Test the disaster recovery plan: Conduct regular drills to ensure effectiveness.

Testing

  1. Establish a comprehensive testing framework: Develop a structured approach to testing.
  2. Conduct regular performance testing: Identify bottlenecks and optimize application performance.
  3. Perform security testing: Identify vulnerabilities and strengthen security measures.
  4. Conduct user acceptance testing (UAT): Ensure applications meet user needs and expectations.

Communication

  1. Develop a communication plan: Outline communication channels and messages during an outage.
  2. Train staff in crisis communication: Equip employees to handle customer inquiries effectively.
  3. Establish customer feedback mechanisms: Gather customer input to improve future responses.

Additional Considerations

While this insight focused on general strategies for safeguarding customer experience, it is important to note that the specific approach will vary depending on the nature of a company's business model, product or service lines, and target audience.

How Can We Help?

Transformidy is available to assist your company in understanding its critical customer experience application/software tools, their impacts during an outage, the approach to take on recovery and management, customer response management, and overall performance measurement.

Contact us or set up a 30 minute complimentary consultation for more information on our services, insights, or showcases. We look forward to hearing from you.

FAQ

What caused the October 2025 AWS outage, and how big was its impact?

AWS traced the October 20, 2025 outage to DNS resolution issues for regional DynamoDB service endpoints in US-EAST-1. AWS reported that the initial DNS issue was mitigated early in the incident, while full service normalization took longer because related subsystems remained impaired. Third-party outage reporting estimated broad downstream impact across thousands of organizations and many countries.

Is this an isolated incident or part of a broader pattern?

Part of a broader pattern. Forrester principal analyst Lee Sustar expects 'at least two major multiday outages in 2026,' attributing the trend to hyperscalers diverting investment away from legacy infrastructure toward GPU-centric data centers for AI workloads while aging infrastructure faces growing complexity. Microsoft's Azure, Microsoft 365, Outlook, and Xbox Live also experienced a major global outage beginning October 28, 2025, roughly a week after the AWS incident.

Why did such a narrow technical failure cause such a large business impact?

Because many unrelated companies and services depend on the same underlying cloud region and infrastructure without necessarily understanding the full extent of their shared dependency. The technical trigger can be narrow while the customer-facing blast radius is broad, which makes concentrated cloud dependency a systemic risk rather than an isolated vendor problem.

What should a business actually do in response to this pattern?

Map which customer-facing functions depend on a single cloud provider or region without a tested failover, and treat that map as a live risk register rather than a one-time audit, since Forrester's own analysis suggests the underlying causes are structural and unlikely to resolve in 2026.