Pricing
Transparent, per-second pricing on the best-available GPU across every provider we aggregate. No minimums, no lock-in, no surprises.
Real prices. No hidden fees. See how GPU.ai stacks up.
| GPU Model | VRAM | GPU.ai | CoreWeave | AWS | Azure | Avg Savings |
|---|---|---|---|---|---|---|
| H200 SXM | 141 GB | $3.81/hr | $6.31/hr | $4.97/hr | $12.99/hr | 53% |
| H100 SXM | 80 GB | $2.64/hr | $6.16/hr | $3.93/hr | $12.29/hr | 65% |
| B200 | 192 GB | $5.29/hr | $8.60/hr | $14.24/hr | $15.00/hr | 58% |
| A100 | 80 GB | $2.00/hr | $2.70/hr | $5.12/hr | $4.10/hr | 48% |
Starting rates across the rest of the fleet. Live rates and stock update continuously.
| Model | VRAM | GPU.ai / hr | Avg savings |
|---|---|---|---|
| H100 PCIe | 80 GB HBM3 | $2.49/hr | 60% |
| A100 40GB | 40 GB HBM2e | $1.45/hr | 45% |
| L40S | 48 GB GDDR6 | $1.10/hr | 40% |
| RTX 6000 Ada | 48 GB GDDR6 | $0.95/hr | 38% |
| A6000 | 48 GB GDDR6 | $0.79/hr | 32% |
| RTX 4090 | 24 GB GDDR6X | $0.44/hr | 35% |
| L4 | 24 GB GDDR6 | $0.42/hr | 33% |
| V100 | 16 GB HBM2 | $0.29/hr | 25% |
Starting rates shown. See live availability →
Metering starts when your instance hits SSH-ready and stops the second you do. No hourly rounding, no idle reservation fees.
Every launch is matched to the cheapest provider that has stock for your exact spec — automatically, on every request.
Move between providers and regions whenever you like. The CLI and API stay identical, so your workflow never changes.
Add credit and we match it 100% — up to $250 in bonus credit across your deposits. Every dollar is refund-guaranteed.
Per second, from the moment your instance reaches SSH-ready to the moment you stop it. No hourly rounding, no minimums, no idle reservation fees.
We continuously scan live capacity across every connected provider and route each launch to the cheapest one that matches your spec — so you get the floor price without checking five dashboards.
No surprise networking fees. You pay the GPU rate; data transfer for normal training and inference workloads is included.
Add credit to your account and we match it 100% — up to $250 in bonus credit across your deposits (manual top-ups and auto-pay both count). Matched credit is promotional and spend-only, and every dollar you spend is refund-guaranteed.
Your GPU is dedicated to you while the instance runs — it's a whole physical GPU, never shared or sliced between tenants. The physical server it sits in, however, is shared with other tenants (each on their own GPU), and every launch is placed on the best available host at that moment. That means a heavy co-tenant can still contend for host resources like system memory bandwidth, PCIe, and networking, so run-to-run performance can vary and we can't promise zero risk of noisy neighbors — even at the same GPU spec, region, and price. Launching during off-peak hours (early morning or late night) tends to land you a cleaner host. For consistent, repeatable performance on recurring workloads, use Reserved, which pins you to dedicated committed capacity.
Serverless scales to zero, so the first request after an idle period can incur a cold start while capacity spins up, and latency can rise briefly under heavy load as we autoscale to meet demand. Steady-state latency is fast; if you need guaranteed low latency at all times, a dedicated on-demand or reserved instance is the better fit.
Yes. Reserved instances trade a commitment for a deeper discount on multi-week and multi-month workloads, and pin you to dedicated capacity for consistent performance. Note that a reservation bills for the committed term whether or not the capacity is actively used — see the Reserved product.