Here's a pricing table that has not changed much all year, and that a surprising number of engineering leaders have never actually computed:
| Type | Typical rate (H-class equivalent) | Availability | Preemption |
|---|---|---|---|
| Hyperscaler on-demand | $3.00–4.00 / GPU-hr | Immediate | None |
| Hyperscaler reserved / committed | $1.70–2.40 / GPU-hr | Contracted | None |
| Neo-cloud reserved | $1.40–2.00 / GPU-hr | Contracted | None |
| Open / decentralized market | $0.90–1.60 / GPU-hr | Best-effort | Possible |
The spread between a committed rate and an open-market rate for comparable hardware is commonly 30–50%. That's a large, persistent number. Persistence like that usually means the market is pricing something real, and it is: certainty. A commit buys you the guarantee that capacity exists when you need it, at a price fixed in advance. An open market gives you the cheapest hours in the industry, but nobody promises they'll be there tomorrow.
The break-even everyone gets wrong
Teams compare the two headline rates and conclude: "committed is cheaper, sign the contract." That comparison is wrong, because a reserved hour is only cheaper if you use it. The right comparison is effective cost per useful hour:
Effective reserved price = reserved rate ÷ realized utilization
Commit at $2.00/hr and actually use 50% of the capacity you're paying for, and your effective price is $4.00/useful hour — worse than just buying on-demand when you need it, and dramatically worse than a market hour you buy only when the job runs. At 90% utilization the same commit delivers $2.22/useful hour, which is genuinely good. The break-even against typical open-market rates lands somewhere in the 45–55% utilization band, and it moves with hardware class, region and how much you value a hard availability guarantee.
Which means the honest question is never "is the committed rate lower?" It's "can I prove — with historical data, not optimism — that I'll keep this cluster above the break-even?" Most teams cannot, and buy anyway. That's not a scandal; it's a budget process that rewards predictable line items over correct ones.
Why the spread persists
Three reasons, none of them mysterious:
- Availability has real value. A trading desk or a latency-critical product genuinely cannot tolerate preemption. For them the open market isn't cheaper, it's unusable — no arbitrage exists for work you're not allowed to run.
- Interruption risk is priced, not hidden. When a marketplace hour is reclaimed and your job dies at hour nine, the retry costs more than the discount saved. Anyone who's been burned once starts treating the discount as fake. Sometimes it is.
- Procurement friction is real. One signature on a one-year commit is tractable; a hundred small market purchases across a shifting set of suppliers requires tooling, accounting and an appetite for variable spend that most finance departments don't have.
The pattern that actually saves money
The teams getting both certainty and the discount do something unglamorous: they partition workloads by interruption tolerance instead of choosing a single sourcing strategy for everything.
| Workload | Right sourcing | Why |
|---|---|---|
| Long, tightly-coupled pretraining run | Committed capacity (your own, or reserved) | Restart cost is brutal; utilization is provably high |
| Serving / inference baseline | Committed, sized to the 20th percentile of load | Steady floor, latency-critical |
| Inference peaks, batch jobs, evals | Open market, bought at run time | Interruption-tolerant, spiky, price-sensitive |
| Hyperparameter sweeps & ablations | Cheapest credible market, checkpointed | Failures are cheap; runs are short |
That partition is the whole trick, and it's the same conclusion we reached from the opposite direction in renting vs buying and, more recently, in the actuarial compute piece: peak-shaped, latency-tolerant work belongs in markets; the provable floor belongs on committed capacity. Teams that adopt it report compute lines 25–45% below their all-committed baseline, without accepting SLA risk they can't tolerate.
The part nobody says out loud
The spread is also a scoreboard on engineering honesty. If your utilization data can't survive a CFO's questions, the committed contract isn't infrastructure — it's a hedge against admitting you don't know your own load. Fix the measurement first. Then the sourcing decision, for most teams, quietly makes itself.
Comments