GPU.ai has a new look. The compute is the same.Read about the redesign  

pricing: live platform rates

The price is the price

On-demand GPUs from every cloud we hold, priced live by the market. Per-second billing, launch in seconds.

Live platform rates

27 GPU types online
GPUVRAMRegionsAvailabilityPrice / GPU
RTX PRO 600096G5 regionshigh availability$1.99/hrlaunch
L40S48G4 regionshigh availability$0.80/hrlaunch
L4048G2 regionshigh availability$0.78/hrlaunch
RTX 6000 ADA48G2 regionshigh availability$0.65/hrlaunch
RTX A600048G3 regionshigh availability$0.40/hrlaunch
RTX 5090community32G8 regionshigh availability$0.51/hrlaunch
A10080G6 regionshigh availability$1.23/hrlaunch
RTX 4090community24G7 regionshigh availability$0.46/hrlaunch
H200 SXM141G6 regionshigh availability$3.59/hrlaunch
H200 SXMcommunity141G5 regionshigh availability$3.59/hrlaunch
H100 PCIE80G2 regionsavailable$1.98/hrlaunch
H100 SXM80G5 regionsavailable$3.49/hrlaunch
L424G4 regionsavailable$0.44/hrlaunch
RTX 409024G3 regionsavailable$0.74/hrlaunch
A100community80G5 regionsavailable$1.06/hrlaunch
RTX A4000community16G6 regionsavailable$0.07/hrlaunch
RTX PRO 450032G2 regionsavailable$0.72/hrlaunch
L40Scommunity48G3 regionsavailable$0.79/hrlaunch
RTX 3090community24G7 regionsavailable$0.15/hrlaunch
RTX 509032G2 regionsavailable$0.99/hrlaunch
RTX PRO 6000community96G2 regionsavailable$1.69/hrlaunch
B200192G3 regionsavailable$6.79/hrlaunch
L4community24G2 regionsrunning low$0.32/hrlaunch
RTX 4000 ADAcommunity20G2 regionsrunning low$0.20/hrlaunch
RTX 2000 ADA16G2 regionsrunning low$0.24/hrlaunch
RTX A500024G2 regionsrunning low$0.27/hrlaunch
RTX 4000 ADA20G2 regionsrunning low$0.28/hrlaunch
A4048G2 regionsrunning low$0.49/hrlaunch
H200 NVL141G1 regionrunning low$3.29/hrlaunch
A100 40GBcommunity40G4 regionsrunning low$0.44/hrlaunch
RTX A6000community48G2 regionsrunning low$0.47/hrlaunch
RTX 309024G1 regionrunning low$0.50/hrlaunch
A100 40GB40G1 regionrunning low$0.89/hrlaunch
H100 NVL94G1 regionrunning low$3.19/hrlaunch
H200 NVLcommunity141G3 regionsrunning low$3.62/hrlaunch
V100community16G2 regionsrunning low$0.10/hrlaunch
B200community192G1 regionrunning low$8.76/hrlaunch
RTX 3080community10G2 regionsrunning low$0.10/hrlaunch
H100 NVLcommunity94G2 regionsrunning low$2.78/hrlaunch
RTX PRO 4500community32G1 regionrunning low$0.33/hrlaunch
RTX 4080community16G1 regionrunning low$0.34/hrlaunch
RTX 5080community16G1 regionrunning low$0.67/hrlaunch
H100 PCIEcommunity80G1 regionrunning low$2.88/hrlaunch
H100 SXMcommunity80G1 regionrunning low$3.54/hrlaunch

lowest live on-demand rate per GPU, refreshed continuously · per-second billing

Transparent rates

The rate you see in the catalog is the rate you're billed: per GPU, per second, with no surprise line items.

Per-second billing

Metering starts when your instance hits SSH-ready and stops the second you do. No hourly rounding, no idle reservation fees.

Best-price routing

Every launch is matched to the cheapest provider with stock for your exact spec. Automatically, on every request.

No lock-in

Move between providers and regions whenever you like. The CLI and API stay identical, so your workflow never changes.

faq

Questions, answered

How is billing calculated?

Per second, from the moment your instance reaches SSH-ready to the moment you stop it. No hourly rounding, no minimums, no idle reservation fees.

How do you beat the providers' own prices?

We continuously scan live capacity across every connected provider and route each launch to the cheapest one that matches your spec, so you get the floor price without checking five dashboards.

Are there egress or networking fees?

No surprise networking fees. You pay the GPU rate; data transfer for normal training and inference workloads is included.

Is there a refund guarantee?

Yes. If you hit a real technical issue (a dead GPU, downtime, a broken instance), open a support ticket and we'll refund the affected usage. No risk in trying us.

Is on-demand performance guaranteed?

Your GPU is dedicated to you while the instance runs. It's a whole physical GPU, never shared or sliced between tenants. The physical server it sits in, however, is shared with other tenants (each on their own GPU), and every launch is placed on the best available host at that moment. That means a heavy co-tenant can still contend for host resources like system memory bandwidth, PCIe, and networking, so run-to-run performance can vary and we can't promise zero risk of noisy neighbors, even at the same GPU spec, region, and price. Launching during off-peak hours (early morning or late night) tends to land you a cleaner host. For consistent, repeatable performance on recurring workloads, use Reserved, which pins you to dedicated committed capacity.

Do serverless requests have consistent latency?

Serverless scales to zero, so the first request after an idle period can incur a cold start while capacity spins up, and latency can rise briefly under heavy load as we autoscale to meet demand. Steady-state latency is fast; if you need guaranteed low latency at all times, a dedicated on-demand or reserved instance is the better fit.

Can I reserve capacity for long runs?

Yes. Reserved instances trade a commitment for a deeper discount on multi-week and multi-month workloads, and pin you to dedicated capacity for consistent performance. Note that a reservation bills for the committed term whether or not the capacity is actively used. See the Reserved product.