Anthropic Is About to Pay $6 Billion to Stop Wasting Its Own GPUs

Anthropic is in talks to acquire Israeli AI startup Decart for about $6 billion — its largest known acquisition. The deal would bring inference-efficiency technology that could make its existing GPUs serve up to 8x more queries, ahead of a possible October IPO.

Aug 13, 2026 - 18:36
0 13

Anthropic Is About to Pay $6 Billion to Stop Wasting Its Own GPUs

Let me tell you something that should have every cloud provider, every GPU broker, and every founder running an AI workload sitting up straight today. Anthropic — the company that's spent the last year making headlines for writing nine-figure checks to landlords and bitcoin miners — is now in talks to spend $6 billion on a three-year-old startup with fewer than a hundred employees and maybe tens of millions of dollars in revenue.

And I'm here to tell you: that's not a crazy price. That's the market waking up to the most expensive secret in the AI industry — most of the silicon we've all been buying is sitting half-idle.

The Deal — What Anthropic Is Actually Buying

Bloomberg reported late Wednesday night that Anthropic PBC is in advanced talks to acquire Decart AI, an Israeli startup based in Tel Aviv, for around $6 billion. Reuters confirmed the talks within hours. Both outlets were careful to say no agreement is final and discussions can still collapse — Anthropic declined to comment, and Decart didn't respond. That's standard for deals at this stage, but it's worth repeating, because this bidding war has already whipsawed once.

Here's the timeline. Three days ago, Calcalist, Israel's leading financial paper, reported Decart was in advanced sale talks at a $6 billion to $7 billion valuation and named SpaceX as the likely buyer. TechTimes went deep on the SpaceX angle Monday — the logic being that Musk's company, which swallowed the Grok developer in February only to watch all 11 original xAI co-founders walk within weeks, needed to rebuild AI talent. Musk publicly denied SpaceX was pursuing it. By Wednesday night, Bloomberg and Reuters had both identified Anthropic as the party in advanced discussions, displacing Amazon and Nebius, the Dutch AI cloud operator that's been building infrastructure in Israel. Nvidia had earlier entered advanced negotiations before a bigger competing bid arrived, per Calcalist.

The price tag is roughly a 50% premium to Decart's last private valuation of $4 billion — set in May, under three months ago, when Radical Ventures led a $300 million Series B with Nvidia, eBay Ventures, Adobe Ventures, and Toyota Ventures in the round, plus angel investors including OpenAI co-founder Andrej Karpathy and former Disney CEO Michael Eisner. A 50% markup in 90 days tells you everything about how fast the bidding environment for this specific kind of talent has tightened.

The Problem Anthropic Is Paying to Solve

You don't spend $6 billion to make your product slightly better. You spend it to fix the number that's keeping your IPO from working. And Anthropic's number is ugly.

The company filed a confidential S-1 with the SEC on June 1, targeting a Nasdaq listing as early as October, with Goldman Sachs, JPMorgan, and Morgan Stanley as lead underwriters. Its most recent valuation is roughly $965 billion, set by a $65 billion Series H that closed in late May. The story it has to sell public investors is that gross margins go from roughly 40% today to 77% by 2028 — which would yield $17 billion in cash flow on about $70 billion in revenue. That's the pitch.

Here's the problem: Anthropic is spending an estimated $19 billion on compute in 2026 alone. That's roughly one dollar of hardware cost for every dollar of revenue it takes in. Its inference costs ran 23% over budget in 2025. The compute cost per dollar of revenue fell from about $0.71 in Q1 to a projected $0.56 in Q2 — real progress, but nowhere near the ratio the 77% margin story requires.

Inference is the whole ballgame. Training happens once; inference happens billions of times a day. Every Claude query that runs is a bill Anthropic pays to its cloud providers and its own infrastructure. You cannot grow your way out of that problem by selling more subscriptions — the cost scales with the usage. You can only fix it by making each query cheaper to serve.

The Secondary Bottleneck Nobody's Talking About — Your GPUs Are Half-Idle

Here's the part that should scare and excite every operator in this industry at the same time. The industry-wide Model FLOPS Utilization — the measure of how much of a chip's theoretical compute actually does productive work — runs between 40% and 50%. Think about that. Half the computing power we've all been paying for, waiting for, and fighting over is sitting idle inside the machine because the software can't wake it up.

Standard AI frameworks like PyTorch and TensorFlow generate GPU execution code with general-purpose compilers that apply the same heuristics to every chip and every model. They don't exploit the specific memory hierarchy, the tensor core scheduling, the parallel architecture of a given design. So a substantial fraction of the chip's raw capacity just idles. This is the dirty secret of the AI buildout: the bottleneck was never just the number of chips. It's how little of each chip we actually use.

Decart's answer is the Decart Optimization Stack — DOS — a vertically integrated layer that sits between models and silicon. It co-designs model architecture with chip constraints, writes hand-optimized kernels for specific chip families, and uses proprietary compilers that generate chip-specific execution code instead of vendor defaults. The results, per the company's May funding announcement: 1,600 tokens per second for agentic inference — roughly eight times the industry average of around 200 — and full-HD video generation at up to 100 frames per second. On Amazon's Trainium chips specifically, Decart says it achieves more than 80% Model FLOPS Utilization. AWS's own documentation validates the work, citing 4x higher frame throughput and 2x better cost efficiency on Trainium versus top GPUs.

And critically for Anthropic — which runs across AWS, Google Cloud, and SpaceX's Colossus facility in Memphis — DOS is hardware-agnostic. It works on Nvidia GPUs, Amazon Trainium, and Google TPUs. A chip-optimization layer that works across every cloud without forcing a hardware monoculture is worth more to a multi-cloud operator than one that locks you into a single vendor.

Here's the math that makes the $6 billion sane: for a company spending $19 billion a year on compute, a technology that more than doubles the productive output of the same hardware is operationally equivalent to buying $10 billion to $15 billion in additional capacity — without buying a single additional GPU. That's not a vanity acquisition. That's a cost-structure acquisition.

Why This Flips the Capex Story

For two years, the entire AI infrastructure narrative has been one word: more. More GPUs. More gigawatts. More data centers. OpenAI is estimated to be spending $50 billion on compute this year and has raised its cumulative target through 2030 to roughly $750 billion. Every earnings call, every hyperscaler announcement, every land purchase and gas plant deal we've covered in this column has been about adding capacity.

This deal is the first loud signal that the frontier labs have hit the other side of that curve. You can't out-buy your way to profitability when the market is about to grade your margin trajectory in public. So Anthropic is doing the thing every good operator eventually does: it's going back to the hardware it already owns and asking how to make it sweat. The efficiency pivot is real, and it's going to change how the whole industry values infrastructure.

The Two Readings — Both True at Once

There are two ways to read a $6 billion check for a 100-person company with tens of millions in revenue, and you need to hold both in your head at the same time.

Reading one: this is a strength signal. Anthropic is engineering-led, and it's buying the kind of capability that produced FlashAttention — the kernel that delivered a 7.6x PyTorch speedup on standard GPUs and quietly made the whole modern LLM era cheaper. Teams like that are how you build a durable cost advantage. If the deal closes, an investor assumption — "inference efficiency will improve" — becomes a confirmed engineering capability, demonstrated at commercial scale, sitting inside the company before the roadshow.

Reading two: this is margin desperation. Anthropic is paying a 50% premium over a valuation set 90 days ago, to a company whose founders have publicly said they want to build something Google-sized — not become an engineering division inside a bigger firm. The talks are early and could collapse; three other bidders were named in the last week alone, and none of them closed. And the whole exercise is happening on a deadline: a deal closed before an October listing carries a completely different narrative weight than one that closes after. When the market's breathing down your neck and your margin story needs proof, you pay up. That's not weakness — but let's not pretend it's pure confidence either.

What This Means for Independent Hosting Providers

If you're running an independent hosting business — and that's who I talk to — this deal has four lessons you can bank.

First — start measuring MFU, not just utilization. If you're renting or owning GPUs, know what percentage of theoretical compute your workloads actually extract. If it's in the 40-50% industry range, you're paying for capacity you're not using. The efficiency vendors are coming for your market, and the operators who understand their own utilization numbers will negotiate from strength.

Second — efficiency software is now a legitimate competitive weapon for smaller operators. The same math that makes DOS valuable to Anthropic applies to you, at smaller scale. If you can serve the same workload on fewer chips, you can undercut every provider who's still pricing on raw scarcity. The next wave of AI infrastructure competition isn't about who buys the most hardware — it's about who wastes the least.

Third — don't build your pricing on GPU scarcity forever. The margin pressure that's driving Anthropic to buy efficiency is the same pressure that will eventually push cloud pricing down as utilization rates climb across the industry. Lock in contracts and position your value prop on service, reliability, and outcome — not on "we have the chips." The chips are about to go further for everyone.

Fourth — understand that talent is what's really being priced. Three Unit 8200 veterans, under a hundred engineers, tens of millions in revenue, $6 billion. That's not a revenue multiple. That's a capability multiple. When the frontier labs start paying capability multiples, every AI engineering hire in the market gets more expensive, and every experienced infrastructure engineer becomes harder to keep. Plan your hiring and retention accordingly.

The Bottom Line

Anthropic just told us where the AI industry is heading, whether this specific deal closes or not. The era of throwing silicon at the problem — buying land, building gigawatts, stacking GPUs like firewood — is entering its endgame. The next competitive frontier is making the hardware we already own work twice as hard.

That's not a bearish story. It's the most bullish efficiency story this industry has produced since FlashAttention. But it's a warning to every operator who built their business model on the assumption that the answer to every problem is another rack of GPUs.

The guys who figure out how to sweat the silicon they already have are going to eat the lunch of the guys still standing in line to buy more. Ent?

— Allan Ali, Founder

This article was produced with AI-assisted research and editorial support. Sources: Bloomberg, Reuters, TechTimes, Calcalist, Haaretz, Decart AI.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Wow Wow 0
Sad Sad 0
Angry Angry 0
Allan Ali

Publisher of Global1.News. Automation architect, systems builder, and the guy making sure the truth gets published.

Comments (0)

User