serverless inference
Every model behind our OpenAI-compatible inference API, priced per use. Point your existing SDK at api.gpu.ai/v1 and pay per token or per image. No instance to rent, no cold infrastructure to babysit.
| Model | Author | Context | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|---|---|
| GLM 5.2 | Z.AI | 198K | $1.40 | $0.260 | $4.40 |
| GLM 5.3 | Z.AI | 198K | $1.40 | $0.260 | $4.40 |
| GLM 5.3 Flash | Z.AI | 1024K | $0.15 | $0.030 | $0.50 |
| GPT-OSS 120B | OpenAI | 128K | $0.15 | — | $0.60 |
| Kimi K3 | Moonshot AI | 1024K | $2.70 | $0.270 | $13.50 |
| Llama 3.3 70B Instruct | Meta | 32K | $1.04 | — | $1.04 |
| Minimax M3 | MiniMax | 512K | $0.30 | $0.060 | $1.20 |
| Muse Glimmer 30B | Meta | 128K | $0.35 | $0.040 | $1.50 |
| Qwen3.5 9B | Qwen | 256K | $0.17 | — | $0.25 |
| Qwen3.6 Plus | Qwen | 977K | $0.50 | — | $3.00 |
| Qwen3.7 Plus | Qwen | 977K | $0.32 | — | $1.28 |
| Qwen3.8 2.4t A95B | Qwen | 986K | $2.00 | $0.250 | $6.00 |
| Qwen3.8 Flash | Qwen | 977K | $0.15 | — | $0.47 |
| Model | Author | Price |
|---|---|---|
| Flash Image 2.5 | $0.039 per image | |
| Flash Image 3.1 | $0.047 per image | |
| Flash Image 3.1 Lite | $0.069 per image | |
| FLUX.1 Kontext Max | Black Forest Labs | $0.080 per output megapixel |
| FLUX.1 Kontext Pro | Black Forest Labs | $0.040 per output megapixel |
| FLUX.1.1 Pro | Black Forest Labs | $0.040 per output megapixel |
| FLUX.2 Dev | Black Forest Labs | $0.015 per image |
| FLUX.2 Flex | Black Forest Labs | $0.030 per image |
| FLUX.2 Max | Black Forest Labs | $0.070 per output megapixel |
| FLUX.2 Pro | Black Forest Labs | $0.030 per image |
| Gemini 3 Pro Image | $0.134 per image | |
| GPT Image 1.5 | OpenAI | $0.034 per image |
| GPT Image 2 | OpenAI | $0.053 per image |
| Qwen Image | Qwen | $0.0058 per image |
| Qwen Image 2.0 | Qwen | $0.035 per image |
| Qwen Image 2.0 Pro | Qwen | $0.075 per image |
| Seedream 3.0 | ByteDance | $0.018 per image |
| Seedream 4.0 | ByteDance | $0.030 per image |
| Seedream 5.0 Lite | ByteDance | $0.035 per image |
| Wan2.6 Image | Wan | $0.030 per image |
| Model | Author | Price |
|---|---|---|
| FLUX 3 | Black Forest Labs | $0.170 per second of video |
| Hailuo 02 | MiniMax | $0.056 per second of video |
| Happyhorse 1.0 T2v | Alibaba | $0.240 per second of video |
| Happyhorse 1.1 T2v | Alibaba | $0.140 per second of video |
| Minimax H3 | MiniMax | $0.139 per second of video |
| Pixverse V5 | PixVerse | $0.060 per second of video |
| Seedance 1.0 Lite | ByteDance | $0.029 per second of video |
| Seedance 1.0 Pro | ByteDance | $0.113 per second of video |
| Seedance 2.0 | ByteDance | $0.160 per second of video |
| Seedance 2.5 | ByteDance | $0.115 per second of video |
| Vidu Q1 | Vidu | $0.044 per second of video |
| Vidu Q3 | Vidu | $0.098 per second of video |
| Vidu Q3 Turbo | Vidu | $0.195 per second of video |
| Wan2.7 T2v | Wan | $0.020 per second of video |
Every model is served through one OpenAI-compatible endpoint at https://api.gpu.ai/v1. Chat models bill per input and output token; the part of a prompt a partner serves from its prompt cache bills at the lower cached-input rate where one is listed (a dash means the model has no cached rate and every prompt token bills at the input rate). Image and video models bill per generation. Prices are the live catalog rates, the same figures the API and the console bill against: no subscriptions, no minimums, billed against your GPU.ai balance. Models marked economy trade a cold start of roughly 30–60 seconds for a lower rate. This page refreshes every few minutes from the live model catalog.