pricing: live platform rates
On-demand GPUs from every cloud we hold, priced live by the market. Per-second billing, launch in seconds.
| GPU | VRAM | Regions | Availability | Price / GPU | |
|---|---|---|---|---|---|
| RTX PRO 6000 | 96G | 5 regions | high availability | $1.99/hr | launch |
| L40S | 48G | 4 regions | high availability | $0.80/hr | launch |
| L40 | 48G | 2 regions | high availability | $0.78/hr | launch |
| RTX 6000 ADA | 48G | 2 regions | high availability | $0.65/hr | launch |
| RTX A6000 | 48G | 3 regions | high availability | $0.40/hr | launch |
| RTX 5090community | 32G | 8 regions | high availability | $0.51/hr | launch |
| A100 | 80G | 6 regions | high availability | $1.23/hr | launch |
| RTX 4090community | 24G | 7 regions | high availability | $0.46/hr | launch |
| H200 SXM | 141G | 6 regions | high availability | $3.59/hr | launch |
| H200 SXMcommunity | 141G | 5 regions | high availability | $3.59/hr | launch |
| H100 PCIE | 80G | 2 regions | available | $1.98/hr | launch |
| H100 SXM | 80G | 5 regions | available | $3.49/hr | launch |
| L4 | 24G | 4 regions | available | $0.44/hr | launch |
| RTX 4090 | 24G | 3 regions | available | $0.74/hr | launch |
| A100community | 80G | 5 regions | available | $1.06/hr | launch |
| RTX A4000community | 16G | 6 regions | available | $0.07/hr | launch |
| RTX PRO 4500 | 32G | 2 regions | available | $0.72/hr | launch |
| L40Scommunity | 48G | 3 regions | available | $0.79/hr | launch |
| RTX 3090community | 24G | 7 regions | available | $0.15/hr | launch |
| RTX 5090 | 32G | 2 regions | available | $0.99/hr | launch |
| RTX PRO 6000community | 96G | 2 regions | available | $1.69/hr | launch |
| B200 | 192G | 3 regions | available | $6.79/hr | launch |
| L4community | 24G | 2 regions | running low | $0.32/hr | launch |
| RTX 4000 ADAcommunity | 20G | 2 regions | running low | $0.20/hr | launch |
| RTX 2000 ADA | 16G | 2 regions | running low | $0.24/hr | launch |
| RTX A5000 | 24G | 2 regions | running low | $0.27/hr | launch |
| RTX 4000 ADA | 20G | 2 regions | running low | $0.28/hr | launch |
| A40 | 48G | 2 regions | running low | $0.49/hr | launch |
| H200 NVL | 141G | 1 region | running low | $3.29/hr | launch |
| A100 40GBcommunity | 40G | 4 regions | running low | $0.44/hr | launch |
| RTX A6000community | 48G | 2 regions | running low | $0.47/hr | launch |
| RTX 3090 | 24G | 1 region | running low | $0.50/hr | launch |
| A100 40GB | 40G | 1 region | running low | $0.89/hr | launch |
| H100 NVL | 94G | 1 region | running low | $3.19/hr | launch |
| H200 NVLcommunity | 141G | 3 regions | running low | $3.62/hr | launch |
| V100community | 16G | 2 regions | running low | $0.10/hr | launch |
| B200community | 192G | 1 region | running low | $8.76/hr | launch |
| RTX 3080community | 10G | 2 regions | running low | $0.10/hr | launch |
| H100 NVLcommunity | 94G | 2 regions | running low | $2.78/hr | launch |
| RTX PRO 4500community | 32G | 1 region | running low | $0.33/hr | launch |
| RTX 4080community | 16G | 1 region | running low | $0.34/hr | launch |
| RTX 5080community | 16G | 1 region | running low | $0.67/hr | launch |
| H100 PCIEcommunity | 80G | 1 region | running low | $2.88/hr | launch |
| H100 SXMcommunity | 80G | 1 region | running low | $3.54/hr | launch |
lowest live on-demand rate per GPU, refreshed continuously · per-second billing
The rate you see in the catalog is the rate you're billed: per GPU, per second, with no surprise line items.
Metering starts when your instance hits SSH-ready and stops the second you do. No hourly rounding, no idle reservation fees.
Every launch is matched to the cheapest provider with stock for your exact spec. Automatically, on every request.
Move between providers and regions whenever you like. The CLI and API stay identical, so your workflow never changes.
faq
Per second, from the moment your instance reaches SSH-ready to the moment you stop it. No hourly rounding, no minimums, no idle reservation fees.
We continuously scan live capacity across every connected provider and route each launch to the cheapest one that matches your spec, so you get the floor price without checking five dashboards.
No surprise networking fees. You pay the GPU rate; data transfer for normal training and inference workloads is included.
Yes. If you hit a real technical issue (a dead GPU, downtime, a broken instance), open a support ticket and we'll refund the affected usage. No risk in trying us.
Your GPU is dedicated to you while the instance runs. It's a whole physical GPU, never shared or sliced between tenants. The physical server it sits in, however, is shared with other tenants (each on their own GPU), and every launch is placed on the best available host at that moment. That means a heavy co-tenant can still contend for host resources like system memory bandwidth, PCIe, and networking, so run-to-run performance can vary and we can't promise zero risk of noisy neighbors, even at the same GPU spec, region, and price. Launching during off-peak hours (early morning or late night) tends to land you a cleaner host. For consistent, repeatable performance on recurring workloads, use Reserved, which pins you to dedicated committed capacity.
Serverless scales to zero, so the first request after an idle period can incur a cold start while capacity spins up, and latency can rise briefly under heavy load as we autoscale to meet demand. Steady-state latency is fast; if you need guaranteed low latency at all times, a dedicated on-demand or reserved instance is the better fit.
Yes. Reserved instances trade a commitment for a deeper discount on multi-week and multi-month workloads, and pin you to dedicated capacity for consistent performance. Note that a reservation bills for the committed term whether or not the capacity is actively used. See the Reserved product.