The Same H100 GPU Rents for 80 Cents or 97 Dollars — Cloud Pricing Is a Joke
The same NVIDIA H100 GPU rents for 80 cents an hour on specialty clouds and $97 at hyperscalers — a 122x spread. Allan breaks down the pricing chaos and what it means for hosting providers.
The Same H100 GPU Rents for 80 Cents or 97 Dollars — Cloud Pricing Is a Joke
Let me tell you something that's been grinding my gears all week. I've been running hosting infrastructure for over a decade, and I've seen some creative pricing in my time — but the cloud GPU market right now is a whole new level of madness. The same NVIDIA H100 — same chip, same memory, same everything — rents for 80 cents an hour on one platform and 97 dollars and change on another. Same silicon. 122 times the price. And the people charging the 97 dollars want you to believe that's normal.
It's not normal. It's not economics. It's the absence of a functioning market, and it's costing real businesses real money every single day. Let me break down what the numbers actually show, because this isn't a story about one bad deal — it's a story about the whole pricing architecture of the AI buildout being held together by opacity.
The 122x Spread — What the Numbers Actually Say
GPU Tracker, a service that scrapes and publishes cloud GPU prices, tracks 5,213 live listings across 54 providers, 75 GPU models, and 130 regions worldwide — the most complete public picture of this market that exists. Here's what it shows: cloud H100 prices range from 80 cents per hour on a spot marketplace to $97.44 per hour on a hyperscaler reserved multi-GPU configuration. That's a 122x spread for the same physical chip.
Now, I know what the apologists will say. "But Allan, reserved multi-GPU bundles include networking, storage, support..." Buh hold on. We're talking about the same H100 — and the median price across all tracked listings is $8.97 an hour, while the 25th-percentile price — what a careful buyer can actually get — is $3.50. The spread isn't a quality gradient. It's a fog of war. The people paying $97 an hour aren't getting 27 times more GPU. They're getting the same chip with a bigger logo on the invoice.
The Hyperscaler Tax — Paying Double for the Same Silicon
Here's the number that should make every finance person in tech sit up straight: hyperscalers charge 99 percent more than specialty cloud providers for the same H100. The median H100 on AWS, GCP, Azure, and Oracle is $13.96 an hour. The median on specialty clouds — CoreWeave, Lambda, Crusoe, RunPod, Nebius, Vast.ai — is $7.00 an hour. Same chip, twice the price, and the only difference is whose name is on the data center.
It gets worse when you look at specific examples. Spheron's July 2026 comparison shows AWS listing H100s at $6.88 an hour while Azure wants $12.29 — for identical hardware. Vast.ai spot listings have gone as low as 34 cents an hour. Thunder Compute offers $1.40 on-demand. The spread isn't between good and bad providers — it's between providers who publish honest prices and providers who charge what the traffic will bear.
Why the Spread Exists — Concentration and the Bundle Shell Game
You want to know why hyperscalers get away with a 99 percent premium? Because they control the listing. GCP alone accounts for 39.8 percent of all publicly tracked cloud GPU listings. The top 10 providers control 93.7 percent of listings, and the big four hyperscalers — AWS, GCP, Azure, Oracle — hold 69.2 percent of publicly listed GPU capacity. When you control seven out of every ten listings, you don't need to compete on price. You just need to make comparison shopping painful enough that nobody does it.
And they've made it very painful. Hyperscaler pricing is buried in instance-type menus, commitment tiers, egress charges, and multi-GPU bundles that make per-unit comparison nearly impossible. That's not an accident — it's a pricing strategy. The bundle is the shell game: hide the per-GPU cost inside a configuration so complex that the buyer's eyes glaze over and they just sign.
The Spot Market Is Eating the Old Model
Here's the part that gives me actual hope. 40.5 percent of all publicly listed GPU capacity is now spot or interruptible — 2,109 spot listings versus 3,104 on-demand. Buyers have gotten comfortable with checkpointing and fault-tolerant training, and the spot market is responding with prices that make the hyperscalers look like a hostage situation.
The L40S is the poster child: a 1,391x spread — from 32 cents an hour spot to $445.25 an hour on a hyperscaler bundle — with spot averaging 82 percent below on-demand. The A100 spot market averages $6.56 an hour versus $17.01 on-demand, and the cheapest A100 spot listing is 8 cents an hour — the most cost-efficient datacenter GPU per token for LLM inference on 7B to 70B workloads. Even H100 spot averages 40 percent below on-demand. Consumer cards are in on it too: median RTX 4090 rentals are 60 cents an hour, delivering near-A100 inference performance on 7B-class models at a quarter of the cost.
What's happening is simple: the specialty clouds and marketplaces are building a transparent, competitive price floor, and the hyperscalers are increasingly selling a convenience premium — sometimes, frankly, a cluelessness premium.
The Secondary Bottleneck Nobody Talks About — Pricing Opacity
Here's what I keep coming back to, and it's the thing the industry doesn't want to discuss. The real bottleneck in the GPU cloud market isn't supply — it's information. When the same chip varies 122x in price across public listings, nobody — not buyers, not analysts, not even the providers themselves — has a reliable signal for what GPU compute is actually worth. That distortion cascades everywhere: into AI startup burn rates, into model pricing, into every colocation contract signed in the dark.
I've written before about how confusing market signals distort the whole infrastructure chain. This is that problem at its purest. A founder with a $50,000 budget can get 100 hours of H100 from one provider or 13,000 hours from another — same money, wildly different outcome — and the only thing separating them is whether they found the right search query.
What This Actually Means for Independent Hosting Providers
If you're running an independent hosting or colo business, this chaos is your opportunity — if you're deliberate about it. Here's what I'd do:
First, publish your prices. The single biggest competitive advantage available right now is transparency. The hyperscalers hide their pricing in bundles. You can win every informed customer in the market just by putting honest numbers on your site and letting the comparison do the selling.
Second, build the spot bridge. If you have spare GPU capacity, list it on the marketplaces — RunPod alone aggregates 679 instances from independent operators, and it's the largest specialty cloud by listing count precisely because it lets small operators monetize idle hardware. Your idle GPUs are an asset, not a sunk cost.
Third, price against the spread, not against your neighbor. When the same chip rents for $3.50 at the 25th percentile, you don't need to undercut the hyperscaler's $97. Get within striking distance of the honest market price, then win on service, support, and uptime — the things marketplaces can't offer.
Fourth, buy smart when you expand. Use the spot market for bursty inference workloads and reserve committed capacity only for your steady baseline. The 40 percent spot discount is real money — treat it as part of your capacity plan, not an afterthought.
The Structural Reality — The Market Is Pricing Itself Out of Trust
Don't expect this to fix itself. The H100 Price Index has been flat since February — prices are stable, which sounds good until you remember that "stable" here means the 122x spread is now a permanent feature, not a transition phase. The hyperscalers have no incentive to publish cleaner pricing while they control 69.2 percent of listings. The specialty clouds have no incentive to raise prices while they're winning share. But the buyers are learning — and that's the one force that actually matters.
Every founder who discovers the 25th-percentile price is a customer the hyperscaler lost. Every startup that builds checkpointing into its training pipeline is a spot-market convert for life. That's how this ends: not with a regulation, not with a pricing-policy change from AWS, but with enough informed buyers that the $97-an-hour H100 becomes what it always was — a museum piece.
The Bottom Line
Here's the truth, plain and simple. The cloud GPU market is the least efficient market I've ever seen in two decades of watching this industry, and that inefficiency is a tax on every business trying to do real AI work. Same chip, 122 times the price, and the only variable is who you're buying from. That's not a market — that's a lottery where the house controls the odds.
But you don't have to play. Publish your prices. Use the spot market. Buy at the 25th percentile. And when a sales rep tries to sell you a $97-an-hour H100, ask them one question: "What exactly am I paying 27 times the market rate for?" Watch them squirm. It's the most fun you'll have all week.
-- Allan Ali, Founder
This article was produced with AI-assisted research and editorial support. Reporting is based on sources cited in the article.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)