AWS Went Down Again Yesterday — and Half the Internet Went With It

Let me tell you something I was watching unfold on Friday morning from my phone while my coffee went cold. Around 6:40 a.m. Eastern Time, the outage trackers started lighting up. Apple Pay stopped working. DoorDash orders froze. Reddit went dark. Hulu streams cut out.

Jul 25, 2026 - 12:12
0 1
AWS Went Down Again Yesterday — and Half the Internet Went With It

AWS Went Down Again Yesterday — and Half the Internet Went With It

Let me tell you something I was watching unfold on Friday morning from my phone while my coffee went cold.

Around 6:40 a.m. Eastern Time, the outage trackers started lighting up. Apple Pay stopped working. DoorDash orders froze. Reddit went dark. Hulu streams cut out. PlayStation Network — 7,000 complaints in minutes, players getting kicked out of games mid-session. The common thread? AWS's US-West-2 region in Oregon. A regional internet connectivity problem inside Amazon's cloud that cascaded into a disruption that took down dozens of major platforms for about two and a half hours.

This was the third significant AWS incident in three months. Not the third outage ever. The third in a single quarter.

And here's what's stuck with me since I read the resolution notice around 9 a.m. — we've been telling businesses for years not to put everything in one cloud. But the message is different this time. Because the data on who's actually listening just landed, and it tells a story that the hyperscalers really, really don't want you to hear.

The Outage — What Actually Happened

The details are straightforward, and that's what makes it worse. AWS confirmed seven impacted services: Direct Connect, Global Accelerator, Internet Connectivity, IoT Core, API Gateway, Elastic Compute Cloud, and Elastic Container Service. That's a connectivity problem at the network layer — not a power failure, not a configuration bug, not a software update gone wrong. A regional internet connectivity issue in a single AWS region that brought seven critical services down simultaneously.

If you run infrastructure, you know what that means. The problem wasn't in one service that could fail over independently. It was in the network fabric underneath all of them. When the network layer goes in AWS, there's no "deploy to another AZ" fix — because the connectivity problem affects everything that talks to that region.

AWS said the disruption was resolved by 9 a.m. ET. Nearly two and a half hours of half the internet being broken because one region in Oregon had a bad morning. And the affected companies — Reddit, DoorDash, Hulu, Sony — could only apologize on social media while they waited for Amazon to fix the problem on their end.

That's not a failure of your multi-AZ architecture. That's a failure of the entire concept of building your business on someone else's network when that network has a single regional failure mode.

The Repatriation Data That Landed at the Perfect Time

Here's where the timing gets interesting. The same week this outage happened, Flexera published updated cloud repatriation numbers: 21% of enterprise workloads have now moved back on-premises or to alternative infrastructure. One in five. That's up from roughly 12% two years ago, and it's accelerating.

And the reasons? A recent Luminix analysis of CIO surveys puts it bluntly: "Hybrid emerges as 2026 norm amid outages and regulations." They specifically call out three events that changed minds: the CrowdStrike global outage, the Azure AD disruption, and — you guessed it — the recurring AWS incidents. Each outage doesn't just inconvenience users. It changes the calculus in boardrooms. A VP of Engineering who spent three years migrating everything to AWS because "we need to be in the cloud" has to explain to the CEO why the company's entire customer-facing product went dark because Oregon had network problems.

The two most-cited examples of successful repatriation — 37signals and GEICO — didn't move everything back. They moved their flattest, most predictable, highest-volume workloads. The kind of workloads that don't need elastic scaling every Tuesday afternoon. The kind that run perfectly well on dedicated hardware. But those are exactly the workloads that generate the most consistent revenue, and the most consistent costs.

When you run the numbers on what 37signals saved by leaving the cloud, and you multiply it by 21% of the enterprise market, the dollars get real fast.

Why This AWS Outage Is Different From the Others

I've written about cloud outages before. I wrote about the CrowdStrike incident last year, about the Azure AD outage, about the March 2026 AWS global disruption. Each time, the hyperscalers issued a post-mortem, added some redundancy, and life went on.

But this one is different for three reasons.

First, the frequency is accelerating. Three significant incidents in three months is not a blip — it's a trend. AWS runs more infrastructure than anyone, and with that scale comes complexity that even Amazon can't fully control. The more they build, the more attack surface for failures. And with $200 billion in 2026 capex driving even more infrastructure deployment, the complexity is compounding faster than the reliability engineering can keep up.

Second, the outage hits at a specific moment in the repatriation cycle. The 21% figure from Flexera means the repatriation trend has passed the "early adopter" phase and is entering the mainstream. Every boardroom that was already considering a hybrid strategy now has fresh evidence — yesterday's evidence — that the single-cloud bet is riskier than it looked three years ago. The marginal cost of that conversation just dropped to zero.

Third — and this is the one nobody's talking about — the hyperscalers are about to report Q2 earnings. Amazon reports next week. Wall Street is already watching AI revenue to capex ratios with a microscope. A major regional outage the week before earnings does not make the "trust us, we know what we're doing" narrative any easier to sell. Especially when Amazon is spending $200 billion this year on infrastructure that just proved it can fail in a way that takes down half the consumer internet.

The Claude Opus 5 Irony — Something Actually Good Happened Friday

I want to mention one other thing that happened on July 24, because the contrast is instructive. Anthropic launched Claude Opus 5 — a model that the company says comes close to the frontier intelligence of Fable 5 at exactly half the price. Same $5 per million input tokens as Opus 4.8. $25 per million output tokens. That's half of Fable 5's pricing.

This launch matters because it's a real product announcement — an AI model that enterprises can actually deploy at scale without the budget conversation becoming a board-level fight. But here's the irony: the same infrastructure that makes models like Opus 5 possible — the hyperscaler data centers, the cloud networks, the AWS backbone — is the same infrastructure that went down yesterday and took Reddit and Hulu with it.

We're building the most advanced AI models in human history on a cloud foundation that still has regional internet connectivity failures. The models are getting smarter. The foundation isn't getting any more reliable.

What This Actually Means for Independent Hosting Providers

First — this is your moment to talk about reliability, not just price. Every time AWS goes down, there's a surge of inbound interest in alternatives. But most hosting providers respond with "we're cheaper." That's the wrong message. The message should be: "We didn't go down. Our network doesn't have a single regional failure point that takes down every customer at once. Here's our SLA from the last three months. Compare it to AWS's SLA." Price matters, but uptime credibility matters more. Lead with the latter.

Second — target the repatriating workloads specifically. The 21% of workloads coming back from the cloud are the predictable, high-volume, margin-stable ones. Database workloads. Application servers that don't need auto-scaling. Static content delivery. These are exactly the workloads that independent hosting providers handle best. Don't try to compete with AWS on elastic Kubernetes clusters or serverless functions. Compete on the workloads they're bad at — the ones that need consistent performance and don't benefit from the cloud premium.

Third — use the outage as an education moment, not a sales pitch. Your customers may not realize how dependent their favorite apps are on AWS. When you explain that Apple Pay, DoorDash, and PlayStation all went down because one region in Oregon had a network issue, you're making a structural argument about concentration risk. That argument works better when it feels like reporting, not selling. Send an email. Write a blog post. Frame it as "here's what happened and why it matters." The sales come later, when they ask "could this happen to us?"

Fourth — lock your own house down. Every hyperscaler outage triggers a wave of DDoS and social engineering attacks targeting their customers' exposed backup infrastructure. If you host repatriating workloads, make sure your onboarding process includes a security posture review. The last thing you want is a customer moving from AWS to you and getting hit with an attack they were protected from inside Amazon's network. Make their migration safer, not riskier.

The Structural Reality — The Cloud Bet Is Being Tested

The easiest thing for a business to do after an outage is to add another redundancy — another AWS region, another cloud provider, another failover plan. That's what the hyperscalers want you to do. Stay in their ecosystem, just spend more to be "more resilient."

The harder thing — and the smarter thing — is to look at your actual workloads and ask: does this need to be in the cloud at all?

21% of enterprise workloads have already answered that question with "no." The Flexera number is from before this outage. After two and a half hours of Apple Pay being down on a Friday morning, that number is going up.

The hyperscalers built a trillion-dollar industry on the promise that their infrastructure was more reliable than anything you could build yourself. Every outage chips away at that promise. Not catastrophically — no single outage kills AWS. But cumulatively, over years, the evidence builds. And when you combine that cumulative evidence with the rapid acceleration of 21% repatriation, the consequences of $200 billion capex spent on infrastructure that can still fail on a regional connectivity issue, and the fact that the most advanced AI models are launching on the same fragile foundation...

Well. The story writes itself.

The Bottom Line

Yesterday's AWS outage was not the biggest cloud failure in history. It wasn't the longest. It wasn't even the most damaging — two and a half hours is survivable for most services.

But it was the third one in three months. And the accumulation is what matters.

The businesses that treat this outage as an isolated incident will be caught by the next one, and the one after that. The businesses that treat it as a signal — a real signal — that the concentration risk in the cloud market is structural, not transient, will be the ones that build infrastructure strategies that actually survive the next decade.

I'm not saying everyone should leave AWS. I'm saying that if you're building your entire business on a single cloud region, and that region went dark for two and a half hours yesterday, and you didn't feel it, and your customers didn't notice — then you got lucky. Don't confuse luck with architecture.

Me? I'm going to check on my own backups. And then I'm going to send an email to my customers explaining what happened, why it matters, and what I'm doing about it. That's not a pitch. It's just what responsible hosting looks like.

— Allan Ali, Founder

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Wow Wow 0
Sad Sad 0
Angry Angry 0
Allan Ali

Publisher of Global1.News. Automation architect, systems builder, and the guy making sure the truth gets published. Health & Science correspondent.

Comments (0)

User