pricing: live platform rates
On-demand GPUs from every cloud we hold, priced live by the market. Per-second billing, launch in seconds.
| GPU | VRAM | Regions | Availability | Price / GPU | |
|---|---|---|---|---|---|
| L40S | 48G | 4 regions | high availability | $0.80/hr | launch |
| RTX PRO 6000 | 96G | 5 regions | high availability | $1.99/hr | launch |
| L40 | 48G | 2 regions | high availability | $0.78/hr | launch |
| RTX 6000 ADA | 48G | 2 regions | high availability | $0.65/hr | launch |
| RTX 5090community | 32G | 8 regions | high availability | $0.34/hr | launch |
| RTX 4090community | 24G | 7 regions | high availability | $0.30/hr | launch |
| RTX A6000 | 48G | 2 regions | high availability | $0.45/hr | launch |
| A100 | 80G | 7 regions | high availability | $1.07/hr | launch |
| RTX PRO 4500 | 32G | 2 regions | high availability | $0.72/hr | launch |
| H100 SXM | 80G | 5 regions | high availability | $3.49/hr | launch |
| A100community | 80G | 5 regions | high availability | $0.88/hr | launch |
| H200 SXM | 141G | 6 regions | high availability | $3.99/hr | launch |
| RTX 3090community | 24G | 5 regions | high availability | $0.14/hr | launch |
| RTX A4000community | 16G | 5 regions | available | $0.08/hr | launch |
| L40Scommunity | 48G | 4 regions | available | $0.79/hr | launch |
| RTX 4090 | 24G | 3 regions | available | $0.74/hr | launch |
| H100 PCIE | 80G | 2 regions | available | $1.98/hr | launch |
| A100 40GBcommunity | 40G | 6 regions | available | $0.56/hr | launch |
| RTX 5090 | 32G | 2 regions | available | $0.99/hr | launch |
| RTX 5080community | 16G | 7 regions | available | $0.16/hr | launch |
| L4 | 24G | 4 regions | available | $0.44/hr | launch |
| H200 NVLcommunity | 141G | 4 regions | available | $3.61/hr | launch |
| RTX 3080community | 10G | 4 regions | available | $0.10/hr | launch |
| A40 | 48G | 3 regions | available | $0.49/hr | launch |
| H100 NVL | 94G | 3 regions | available | $3.19/hr | launch |
| B200 | 192G | 3 regions | available | $6.79/hr | launch |
| B300 | 288G | 2 regions | available | $7.89/hr | launch |
| H100 PCIEcommunity | 80G | 2 regions | available | $2.54/hr | launch |
| B200community | 192G | 2 regions | available | $5.32/hr | launch |
| RTX A5000community | 24G | 3 regions | running low | $0.23/hr | launch |
| RTX 2000 ADA | 16G | 2 regions | running low | $0.24/hr | launch |
| RTX 6000 ADAcommunity | 48G | 1 region | running low | $0.74/hr | launch |
| H200 SXMcommunity | 141G | 3 regions | running low | $3.98/hr | launch |
| H100 SXMcommunity | 80G | 3 regions | running low | $2.27/hr | launch |
| RTX A6000community | 48G | 2 regions | running low | $0.41/hr | launch |
| RTX 4080community | 16G | 3 regions | running low | $0.18/hr | launch |
| L4community | 24G | 1 region | running low | $0.27/hr | launch |
| RTX A5000 | 24G | 1 region | running low | $0.27/hr | launch |
| RTX PRO 6000community | 96G | 1 region | running low | $1.69/hr | launch |
| A100 40GB | 40G | 1 region | running low | $0.89/hr | launch |
| V100community | 16G | 2 regions | running low | $0.09/hr | launch |
| H100 NVLcommunity | 94G | 2 regions | running low | $2.37/hr | launch |
lowest live on-demand rate per GPU, refreshed continuously · per-second billing
The rate you see in the catalog is the rate you're billed: per GPU, per second, with no surprise line items.
Metering starts when your instance hits SSH-ready and stops the second you do. No hourly rounding, no idle reservation fees.
Every launch is matched to the cheapest provider with stock for your exact spec. Automatically, on every request.
Move between providers and regions whenever you like. The CLI and API stay identical, so your workflow never changes.
faq
Per second, from the moment your instance reaches SSH-ready to the moment you stop it. No hourly rounding, no minimums, no idle reservation fees.
We continuously scan live capacity across every connected provider and route each launch to the cheapest one that matches your spec, so you get the floor price without checking five dashboards.
No surprise networking fees. You pay the GPU rate; data transfer for normal training and inference workloads is included.
Add credit to your account and we match it 100%, up to $150 in bonus credit across your deposits (card top-ups, auto-pay, and crypto deposits all count). Matched credit is promotional and spend-only, and every dollar you spend is refund-guaranteed.
Your GPU is dedicated to you while the instance runs. It's a whole physical GPU, never shared or sliced between tenants. The physical server it sits in, however, is shared with other tenants (each on their own GPU), and every launch is placed on the best available host at that moment. That means a heavy co-tenant can still contend for host resources like system memory bandwidth, PCIe, and networking, so run-to-run performance can vary and we can't promise zero risk of noisy neighbors, even at the same GPU spec, region, and price. Launching during off-peak hours (early morning or late night) tends to land you a cleaner host. For consistent, repeatable performance on recurring workloads, use Reserved, which pins you to dedicated committed capacity.
Serverless scales to zero, so the first request after an idle period can incur a cold start while capacity spins up, and latency can rise briefly under heavy load as we autoscale to meet demand. Steady-state latency is fast; if you need guaranteed low latency at all times, a dedicated on-demand or reserved instance is the better fit.
Yes. Reserved instances trade a commitment for a deeper discount on multi-week and multi-month workloads, and pin you to dedicated capacity for consistent performance. Note that a reservation bills for the committed term whether or not the capacity is actively used. See the Reserved product.