AI Is Breaking Its Own Data Centers - and the Fix Isn't More Chips
Bloomberg's investigation finds AI's volatile power demand is damaging its own data centers - batteries, generators and cooling systems failing early. A founder's take on the hidden reliability risk in the AI buildout.
AI Is Breaking Its Own Data Centers — and the Fix Isn't More Chips
Let me tell you something that's been sitting wrong with me since I read it this morning. I've spent more nights than I care to count staring at power feeds, UPS alarms, and generator test logs. And I can tell you exactly which phone call I dread most: the one where a client says the power equipment is dying and the hardware is fine. That call used to be rare. After what Bloomberg reported today, I think it's about to become the norm.
The headline is brutal and simple: AI's volatile power demand is damaging its own data centers. Not the grid. Not the neighbors. The facilities themselves. Batteries, generators, cooling systems — the boring, unglamorous equipment that keeps the lights on — are malfunctioning and dying years before they should. And the people who run these buildings are quietly watching their most expensive assets eat themselves from the inside.
The Bloomberg Investigation — What They Actually Found
Bloomberg talked to more than three dozen power experts across the US and Europe — generators, data center developers, grid operators, utilities, insurers, regulators. Almost all of them said the same thing: the physical stress on these facilities is real, and it's visible. At xAI's Colossus facility in Memphis, gas-fired turbines developed cracks. At smaller data centers in the UK, the same thing. Cranks on small natural gas combustion engines used for backup power have literally broken off.
The culprit isn't the amount of power. It's the shape of the demand. AI workloads don't draw steady, predictable power like a traditional server room. When you train a model, hundreds of thousands of GPUs power up and down in unison — on a millisecond basis. Amber Villegas-Williamson, a principal consultant at the Uptime Institute, put it the way I would: "It's like over-revving your car wears out the engine faster than keeping a constant speed."
Shannon Miller of Mainspring Energy framed it even better. A gigawatt data center, he said, is equivalent to a city the size of Boston — half of which can flicker on and off every few seconds. And some of the AI campuses planned in Texas and the Midwest are five times bigger than that. We're building city-sized electrical loads that behave like a strobe light, and then we're surprised when the equipment attached to them gives up.
The Numbers Nobody in This Industry Wants to Talk About
Here's where it gets genuinely scary. Drew Baglino, the former Tesla executive who now runs Heron Power Electronics, told Bloomberg that AI power usage can spike as much as 50% above design capacity. A 1 gigawatt facility can pull 1.5 gigawatts for a split second. And that spike isn't gradual — it's instant. Jon Parrella, CEO of energy-storage developer Terraflow, compared it to driving a Ferrari and shifting straight from sixth gear to first. "You can't swing that fast," he said.
Most power equipment isn't designed for that. Batteries that were installed specifically to smooth out these swings have needed replacement within months — sometimes weeks. The Uptime Institute confirmed it. This isn't a Texas problem or a Tennessee problem. It's happening in the Middle East, Africa, Europe, and the US. Everywhere AI compute is being stood up fast, the power gear is being chewed up faster.
And here's the part that should terrify anyone financing these projects: some facilities are seeing actual uptime closer to 80%, not the 99.999% they were financed on. A person involved in financing these facilities told Bloomberg that unless this gets resolved, it could hit investors in certain projects within the next 12 to 24 months. Think about that. The whole AI buildout is priced on the assumption that once these places go online, they run 24/7/365. The reality is some of them are running four days out of five.
The Secondary Bottleneck — The Depreciation Bomb Nobody Modeled
Here's the part that keeps me up at night, and it's the same pattern I've been screaming about for weeks. Everyone models the GPU cost. Nobody models what happens to everything around the GPUs.
Switch's chief strategy officer Jason Hoffman nailed it: "The financial consequence is not primarily replacing a pump or a breaker or some power component — it's the value of that expensive compute capacity not generating revenue because it's offline." Downtime costs at these facilities range from thousands to hundreds of thousands of dollars per minute, depending on the workload. You don't lose a pump. You lose a million dollars an hour of Nvidia hardware sitting dark.
And it's not just uptime. The rate of depreciation on GPU racks themselves has already raised questions about whether this industry can be as profitable as it promises. Now add premature failure of the entire power delivery chain — turbines, batteries, capacitors, transformers, cooling — and the balance sheet math gets worse in a hurry. The industry is already borrowing hundreds of billions against these assets. Those assets are now wearing out faster than the models assume. That's not an engineering footnote. That's a credit event waiting to happen.
The Grid Risk — This Stops Being Their Problem Fast
Now the part that should worry every one of us, even if you've never touched a data center. This volatility doesn't stay inside the fence line.
Sreemant Roy, a power-quality expert at Schneider Electric, told Bloomberg these loads are "extremely dynamic or fluctuating, which causes grid instability and can lead to, if not corrected, potential blackouts or power outages." Worse, data centers can cause sub-synchronous oscillations in the power flow — a technical way of saying they can damage equipment connected to OTHER parts of the network. Roy said it plainly: "That has made utility companies globally very worried."
The North American Electric Reliability Corporation has spent the last two years warning that data centers are one of the greatest risks to grid stability. It evaluated more than 33 gigawatts of operational data centers and found about three-quarters of their load models are insufficient to represent data-center dynamic behavior. Earlier this year, NERC issued a rare level-three alert requiring big data centers to address these risks — with responses due by August 3. That deadline just passed. We'll see what comes of it.
The Counter-Argument — "We're On It" Doesn't Fix Physics
The industry's response, as always, is that they're aware and working on it. Nvidia says it started working more closely with power experts when it developed Blackwell. Some operators are running "dummy math" — side computations that aren't part of training — just to keep GPU loads steady. The Department of Energy set up a test bed at the National Laboratory of the Rockies so developers can test their gear against AI variability.
All fine. None of it fixes the problem for the buildings that already exist.
The dummy math trick is the most revealing one. Think about what they're admitting: they're burning electricity on purpose — wasting power in the middle of a power shortage — because their own hardware can't handle the swings. That's not a solution. That's a coping mechanism. And it tells you everything about how unprepared this buildout was for the physical reality of AI workloads.
The counter-argument also ignores the retrofit problem. Batteries, capacitors, flywheels, better power conditioning — that equipment exists, but it costs money and time, and the whole industry is in a race to stand capacity up as fast as possible. Nobody stops a $10 billion campus build to install a few million dollars of power conditioning. So they don't. And then the turbines crack.
What This Actually Means for Independent Hosting Providers
First — if you're running a colocation or hosting operation, start treating power quality as a first-class monitoring metric, not a footnote. Watch your UPS battery health like a hawk. AI neighbors on shared infrastructure can send voltage transients your way, and your gear doesn't care where they came from.
Second — read your power contracts and your uptime SLAs with fresh eyes. If your facility is seeing 80% uptime instead of 99.999%, the revenue loss is on you, not the utility. Build your pricing models around the worst case, not the brochure case.
Third — if you're buying new capacity, favor facilities with serious power-conditioning infrastructure — flywheels, capacitors, proper harmonic filtering — even if it costs more per kilowatt. The cheapest power today is the most expensive power next year when the batteries die in month three.
Fourth — watch the depreciation math in your own business. If the hyperscalers are discovering their power gear wears out faster than modeled, your own UPS and generator replacement cycles are probably optimistic too. Budget for it now.
The Structural Reality — The Buildout Just Got More Expensive
Here's the uncomfortable truth: this story doesn't kill the AI buildout. It makes it more expensive. The West Texas campus Joulent is developing with Chevron — 2.67 gigawatts, built to Microsoft's 99.999% reliability spec — already pushed power delivery from 2027 to 2028 just to engineer around the volatility problem. When projects start slipping a year because of power quality, the cost of AI compute goes up, the timeline stretches, and the returns shrink.
That's the pattern I've been pointing at for weeks now: every bottleneck in this buildout — chips, cooling, water, community consent, debt markets — gets solved by spending more money and waiting longer. This one is no different. The difference is this bottleneck is invisible until it breaks. And it's breaking right now, in the buildings already running, at the exact moment the industry is trying to convince lenders it's all under control.
The Bottom Line
I've said it before and I'll say it again: the AI buildout isn't a software story. It's a physics story wearing a software costume. The GPU gets the headline. The transformer gets the blame. And the batteries — the unsung workhorses that take the hit every time a training run spikes — they just die quietly, weeks ahead of schedule, taking your uptime and your margins with them.
If you're a founder running infrastructure, this is your warning shot. The hyperscalers can absorb a few cracked turbines. Can you? The answer determines whether you're still standing when this shakeout hits — or whether you're the one explaining to investors why the equipment failed before the loan did. Plan for it now, ent? The physics isn't going to wait for the next earnings call.
— Allan Ali, Founder
This article was produced with AI-assisted research and editorial support. Reporting is based on sources cited in the article.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)