Every few weeks someone posts a thread doing FLOPs math on a training run. Multiply parameters by tokens by six, divide by accelerator FLOPS, announce a number, done. This math is never wrong and always useless, because FLOPs-to-dollars is the least interesting conversion in the industry. The interesting conversions happen after the math stops.
Let's do the full stack for two realistic scenarios: a serious fine-tune of a 13B model, and a from-scratch pretraining run at 7B scale. Prices are as of this writing; treat them as a snapshot, not scripture.
Scenario one: the serious fine-tune
Say you're fine-tuning a 13B model on 2B tokens with LoRA at moderate rank. On paper this is maybe 30,000 H100-hours. Nobody's budget survives contact with that number, so let's be honest about the multiplier.
1. Utilization, or why you pay for 30,000 hours and use 19,000
Real-world MFU (model FLOPs utilization) on a well-tuned job lands around 35-45%. Bad data pipelines, checkpoint stalls and straggler nodes eat the rest. You pay for allocated hours, not useful ones. Budget a 1.5-1.8x multiplier over ideal math.
2. Failed runs
Hyperparameter sweeps, OOM crashes, a corrupted shard at hour nine, the run you killed at hour twelve because the loss curve looked wrong. Mature teams budget 20-30% waste here. First-time teams discover it.
3. The blob tax
Checkpoints. A 13B model with optimizer states is ~150GB per full save. Save every 500 steps across a multi-day run and you'll push tens of terabytes. Storage is cheap; egress is where hyperscalers get you — $0.05-0.12/GB at list rates adds up fast if your storage bucket and your cluster live in different places.
4. The humans
Two ML engineers for three weeks. If your fully-loaded cost per engineer-month is $40K, that's $60K of salary attached to a compute bill — often the largest single line item, and the one nobody's pricing page mentions.
Running the numbers
| Line item | Hyperscaler (on-demand) | Decentralized market |
|---|---|---|
| 30K ideal hrs × 1.6 utilization | 48,000 hrs × $3.00 = $144,000 | 48,000 hrs × $1.80 = $86,400 |
| Failed-run waste (+25%) | $36,000 | $21,600 |
| Storage + egress | $4,000-8,000 | $2,000-4,000 |
| Human time | $60,000 | $60,000 (retries cost wall-clock too) |
| Realistic total | ~$250,000 | ~$172,000 |
Two honest caveats. First, a hyperscaler with a committed discount closes much of the GPU-rate gap — if you can get one, and if your job is big enough to be worth negotiating. Second, decentralized markets charge you in engineering time: expect to spend a day or two extra on tooling, checkpoint-retry logic and supplier vetting. The table above prices compute; your patience is a real cost too.
Scenario two: pretraining at 7B
Pretraining changes the shape of the problem. You're now at 100K-1M GPU-hours, utilization matters more (well-run jobs hit 45-55% MFU), and the egress/storage tax shrinks relative to the total. The failure multiplier matters less in percentage terms but more in absolute ones — a crashed week-three run on 4,000 GPUs is a five-figure mistake you make exactly once.
At this scale, most teams end up in one of three buckets: hyperscaler commits (expensive, reliable, negotiable), neo-clouds (10-20% cheaper, less slack), or hybrid setups — pretraining on committed capacity, with decentralized markets soaking up sweeps, ablations and the long tail of experiments. That hybrid pattern is, in our view, where 2026 budgets are actually going. The experiments you'd never run because they cost $8,000 suddenly cost $3,800, and you run twelve of them instead of five.
What we'd tell a team starting today
- Price your run in allocated hours, not ideal FLOPs, and multiply by 1.7 before you fall in love with the project.
- Put egress and storage in the budget on day one. It's small until it isn't.
- Match the workload to the market. Pretraining core = committed capacity. Sweeps, fine-tunes, RLHF rollouts, evals = cheapest credible market you can verify.
- Ask any decentralized provider how they verify compute before you ask about price. Verification is the whole ballgame.
The brochure rate is a marketing number. The invoice is the truth. Budget for the invoice.
Comments