OpenAI Just Gave Away Its Best Model for Free — and That's the Most Expensive Thing in Tech

OpenAI removed text chat limits for all ChatGPT users, making GPT-5.6 Luna the default for a billion weekly users and cutting API prices by 80 percent. Token prices fell from two dollars to one dollar twenty per million since June, and the price war is reshaping AI's trillion-dollar infrastructur...

Aug 09, 2026 - 20:36
0 10

OpenAI Just Gave Away Its Best Model for Free — and That's the Most Expensive Thing in Tech

Let me tell you something that's been sitting with me since the news crossed my desk this week. I've been running infrastructure and watching the AI business from the inside for over a decade, and I've seen plenty of "this changes everything" moments. Most of them were marketing. This one isn't.

On Thursday, OpenAI did something no company with a trillion-dollar valuation has ever done: it made its current-generation model free and unlimited for everyone. Not a free trial. Not a teaser. GPT-5.6 Luna — the same model it was charging developers real money for two weeks ago — is now the default for every free and Go user of ChatGPT, and the text-chat limits are gone. The same week, it cut the API price of Luna by 80 percent. And the numbers behind that decision should terrify every AI CEO on the planet, because the unit price of intelligence is collapsing toward zero — and the entire trillion-dollar buildout is priced on the assumption that it wouldn't.

The Announcement Nobody Wrapped Their Head Around

Let me lay out what OpenAI actually did on August 6, because the coverage buried the lede. ChatGPT crossed 1 billion weekly active users — a number that would make any social network jealous. And on the same day, OpenAI announced that free and Go users get unlimited text chats, powered by GPT-5.6 Luna as their new default model, replacing GPT-5.5. A "Think" button for harder questions lands next week.

For Plus and Pro subscribers, the update brings an upgraded GPT-5.6 Sol tuned for quick tasks — questions, web research, planning, writing — plus a thinking slider that lets you dial how much effort the model puts into an answer. OpenAI's internal evaluation says factual errors are 62 percent less common with Luna and 68 percent less common with Sol compared to GPT-5.5 Instant. The free tier still has separate limits for files, images, voice, and image generation — so this isn't pure charity — but the core product, the thing people actually use all day, just went unlimited at the frontier tier.

Now here's the part that should make every founder sit up: this wasn't a standalone consumer move. Two weeks before the free-tier announcement, OpenAI cut the API price of GPT-5.6 Luna by 80 percent — from $1.00 to $0.20 per million input tokens, and from $6.00 to $1.20 per million output. It cut the mid-tier Terra model 20 percent, to $2 and $12. And it added a "Fast" mode to Sol that runs 2.5 times faster at twice the price. Same intelligence, cheaper, faster, free. That's not a promotion. That's a strategic decision about where the market is going.

The Numbers That Should Terrify Every AI CEO

Here's the data point that actually matters, and almost nobody is talking about it. Silicon Data's LLM Token Expenditure Index — which tracks both the closed frontier providers and the open-weight platforms — shows the average price of a million tokens fell from above $2 at the start of June to $1.20 this week. Two dollars to a dollar twenty in ten weeks. A 40 percent collapse in the unit price of intelligence, with no sign of stopping.

The reason is no mystery: Chinese open-weight models. Moonshot's Kimi K3 — a 2.8-trillion-parameter model that anyone can download and run — spooked Wall Street in July when it proved you don't need a hyperscaler's data center to get frontier performance. Alibaba's Qwen line, DeepSeek's open weights, and a dozen other Chinese labs are doing the same thing: shipping models that are free to download and cheap to run, which forces every closed provider to defend market share with price cuts. OpenAI's 80 percent Luna cut isn't generosity. It's a response to a structural threat.

And the market has noticed. The AI stock sell-off last month was driven by exactly this fear: if the unit price of intelligence keeps collapsing, the revenue that's supposed to pay for a trillion dollars of data centers never materializes. Barclays, Nomura, and Morgan Stanley are all tracking this repricing cycle and what it does to enterprise AI spending. The hyperscalers — Microsoft, Google, Amazon — are the most exposed, because their entire capex thesis assumes sustained high margins on inference.

The Jevons Trap — Why Cheaper AI Doesn't Mean Less Compute

But here's where the narrative splits, and this is the part I want every independent operator to understand. There are two ways to read this price collapse, and they lead to completely different business decisions.

Reading one: the bears are right, and the AI capex bubble pops because nobody can charge enough for tokens to pay for the machines. Reading two: this is the Jevons paradox in real time — when the price of something collapses, usage explodes, and total spending goes up, not down. Steam engines got more efficient, so coal consumption rose. When search went free, server demand exploded. When cloud computing got cheap, the hyperscalers got rich.

The analysts quoted in the South China Morning Post this week are firmly in the second camp. "Competition is up and prices are down," Silicon Data wrote on X. "This is good for consumer and enterprise users of AI and promotes much wider and faster AI adoption." Cheaper tokens mean more agents, more automation, more workloads, more inference requests, more training runs — and every one of those needs physical infrastructure. A billion weekly users, each now free to chat unlimited, is a lot of GPUs humming somewhere. OpenAI just bet its entire infrastructure strategy on usage exploding fast enough to outrun the price collapse.

What This Actually Means for Independent Hosting Providers

If you run a hosting business — and that's who I'm talking to — here's what I'd do with this news, and it's not what most people will tell you.

First, stop pricing your services as if model margins will ever recover. The unit price of intelligence is not going back up. Open weights can't be un-released, Chinese labs will keep undercutting, and every frontier provider will keep matching. Build your business on the workload layer, not the token layer.

Second, treat open-weight models as your friend, not your enemy. Every company that runs a 2.8-trillion-parameter Kimi K3 or a DeepSeek on its own hardware needs a place to put it, power to feed it, and storage for the data. That's your market. The model layer is becoming a commodity; the infrastructure layer is where the margin is moving.

Third, watch utilization, not token prices. If your customers' GPU utilization is climbing, you're winning — because falling token prices are translating into more workloads, not fewer. The winners in this cycle will be the operators who capture the demand explosion, not the ones who try to hold the line on pricing.

Fourth, don't sign long-term contracts that assume you can keep charging AI premiums. The premium is eroding. Build flexible pricing that scales with usage, and be ready to serve customers who suddenly need ten times the compute because their token costs just dropped 80 percent.

The Structural Reality — Zero Is a Destination, Not a Stop

Let me be straight with you about where this ends, because I don't think it ends where the bears or the bulls say it does. The price of intelligence is heading to zero as a practical matter — not because intelligence is worthless, but because competition has made it abundant. That's the whole open-weight thesis: when the best model is downloadable for free, the only thing you can charge for is the ability to run it at scale, reliably, with power and cooling and support.

That's not a crash scenario. That's a margin-migration scenario. The model makers will fight over a shrinking pie of token revenue while the infrastructure providers — the people who actually own the machines, the power, the cooling, the connectivity — watch the volume explode. OpenAI's free tier is a bet that the demand explosion is real. If it's right, the capex stays justified. If it's wrong, the model layer was never the business anyway.

The Bottom Line

So here's my take, and I'm not hedging: the most expensive thing in tech right now is a free product. OpenAI just gave away the thing it spent billions training, to a billion people, while the industry spends a trillion dollars on the machines to serve them. That's either the dumbest business decision in history or the smartest demand-creation play ever executed — and the difference between the two will show up in GPU utilization numbers over the next two quarters, not in press releases.

Me? I'm betting on the machines. The model layer is becoming a commodity, and the infrastructure layer is where the real business has always been. If you're an independent operator, that's your moment. Position for the volume, not the margin. Buh trust me on that one.

— Allan Ali, Founder

This article was produced with AI-assisted research and editorial support. Sources: TechCrunch (Aug 6), The Verge (Aug 6), South China Morning Post (Aug 9), CNBC via wccftech (Jul 30), Silicon Data LLM Token Expenditure Index.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Wow Wow 0
Sad Sad 0
Angry Angry 0
Allan Ali

Publisher of Global1.News. Automation architect, systems builder, and the guy making sure the truth gets published.

Comments (0)

User