Key takeaways
- GPU cloud pricing meters per GPU-hour, with separate line items for data egress and storage that quietly inflate the bill.
- On-demand H100 rentals run roughly $2.99 to $10.98 per GPU-hour as of September 2026, depending on provider. Reserved and spot cost less but trade flexibility.
- The bill is large, spiky, and in US dollars, so a card that gets declined on a spike or adds FX is a real operational risk.
You kicked off a training run on Friday. It was meant to take twelve hours. A checkpoint failed, the job restarted, and it ran through the weekend across more GPUs than you meant to spin up. Monday morning the invoice is five figures. Or an inference endpoint went viral, the fleet autoscaled, and the meter ran the whole time. The per-hour rate looked cheap. The total did not. Understanding how GPU cloud pricing works, and paying for it cleanly, is the difference between a predictable cost and a monthly surprise.
How does GPU cloud pricing work?
GPU cloud compute is priced per GPU-hour: you pay for each hour a GPU runs, plus separate charges for data egress and storage. The headline number you compare across providers is the per-GPU-hour rate. It is only part of the bill. Moving data out of the provider (egress) is metered per gigabyte, and storing datasets, checkpoints, and model weights is billed per gigabyte per month. On a large training job, egress and storage can add a meaningful slice on top of the compute line, so GPU cloud pricing is really three meters, not one.
The per-GPU-hour rate itself comes in three buying modes, and choosing the wrong one for the workload is where teams overpay.
| Mode | What it is | Trade-off |
|---|---|---|
| On-demand | Pay-as-you-go per GPU-hour, no commitment, start and stop anytime. | Highest rate per hour, but full flexibility and nothing owed when idle. |
| Reserved / committed | Commit to a term or a fixed amount of capacity in exchange for a lower rate. | Cheaper per hour, but you pay for the commitment whether you use it or not. |
| Spot | Bid on the provider's spare capacity at a steep discount. | Cheapest, but the provider can reclaim it with little notice, so it fits interruptible jobs only. |
This three-mode split is the part of GPU cloud pricing you actually control. The pattern is consistent across providers. Flexibility costs more. On-demand is the safe default for bursty or experimental work. Reserved makes sense once you know your baseline load and can commit to it. Spot is for training that checkpoints well and can survive an interruption.
How much does an H100 cost per hour?
As of September 2026, on-demand NVIDIA H100 rentals run roughly $2.99 to $10.98 per GPU-hour. The lower end tends to sit at specialist GPU clouds, and the higher end at the large hyperscalers. Some providers advertise even lower spot rates for interruptible capacity. Prices vary by provider, region, and contract, and they move fast, so treat any number as a snapshot and re-check a live source before you budget. See the Thunder Compute H100 pricing breakdown for a current comparison.
The spread matters, and it is the widest single variable in GPU cloud pricing. At the low end, eight H100s for a week of on-demand training is a few thousand dollars. At the high end, the same run is several times that. Reserved and spot pull the effective rate down further, but only if the workload fits the mode.
Why do cloud bills spike?
The per-GPU-hour model is easy to reason about and hard to forecast, and it is the part of GPU cloud pricing that turns a cheap sticker rate into an unpredictable bill. A few things stack up at once.
Runs go long. A failed checkpoint or a sweep that expands mid-week keeps GPUs metering past your estimate. Fleets scale wide. Autoscaling inference adds instances when traffic climbs, and the meter follows. GPUs sit idle. A notebook left running overnight still bills by the hour. And the side charges add up: egress from pulling datasets and storage for months of checkpoints land on the same invoice as compute.
None of this shows up in the per-hour rate you compared when you chose a provider. You are not paying a subscription. You are paying for exactly what ran, and what ran is often more than you planned. Keeping that under control is the job of cloud cost management.
How should a team pay a large, spiky cloud bill?
Cloud invoices are large, arrive on a variable schedule, and are almost always denominated in US dollars. That combination breaks the payment method most teams reach for first. GPU cloud pricing sets the number; how you pay decides whether it clears.
A consumer card has a limit that a five-figure spike can blow through, and a large, unusual charge is exactly what an issuer flags and declines. A decline on your GPU provider is not a billing annoyance. It can suspend the account and interrupt production. If your card settles in another currency, you also pay a foreign-exchange margin on every dollar of every invoice, on a bill that is already your biggest line item.
The fix, the same one behind how AI companies pay for all their tools, is a funding setup built for the shape of the bill. You want a card that carries no FX surcharge on a US-dollar invoice, that you can size per vendor so one provider cannot drain everything, and that you can freeze the moment a job runs away.
What an H100 costs, and the three ways to buy
H100 on-demand
$2.99-$10.98 / hrPer GPU-hour, Sep 2026, varies by provider.
On-demand
Most flexiblePay by the hour, no commitment, highest rate.
Reserved
CheaperCommit for a term for a lower rate, less flexible.
Spot
CheapestSpare capacity at a discount, can be reclaimed.
Prices vary by provider and move fast. Storage and data egress add to the bill. Verify at time of use.
How Endl fits
GPU cloud pricing is only half the equation; paying the bill cleanly is the other half. Endl lets you fund and pay a US-dollar cloud bill without an FX surcharge or a surprise decline.
Card spend in USD carries $0 in added FX. For other currencies you pay the Visa rate plus a flat 1%, with no hidden markup, so the cost is a number you can predict rather than a margin buried in the rate. You issue a virtual card per provider, set a per-card limit that matches that vendor's expected bill, and freeze any card instantly if a run goes sideways. Physical cards are available too, and every card is a debit or spend card, not credit, so you pay from a balance you fund rather than borrowing.
That balance is self-custodial and funded on stablecoin rails at a flat 0.5%, and it settles in under five minutes, 24/7. When an invoice lands off-cycle, you top up and pay the same day instead of waiting on a bank window. You fund before you spend, so the money has to be there. The upside is control: a per-card limit means a runaway job cannot spend past what you set.
Two things to be clear about. Endl is not a bank and balances are not insured. Endl operates as a registered VASP in the EU and an MSB in Canada. Read the pricing page for the full fee schedule, or start free and issue your first per-vendor card.




