Serverless inference
Every model behind our OpenAI-compatible inference API, priced per use. Point your existing SDK at api.gpu.ai/v1 and pay per token or per image — no instance to rent, no cold infrastructure to babysit.
| Model | Author | Context | Price |
|---|---|---|---|
| DeepSeek V4 Flash 0731 | DeepSeek | 1024K | $0.28 / 1M out |
| DeepSeek V4 Pro | Deepseek | 128K | $3.48 / 1M out |
| Gemma 3n E4b It | 32K | $0.12 / 1M out | |
| Gemma 4 31B It | 256K | $0.97 / 1M out | |
| GPT-OSS 120B | OpenAI | 128K | $0.60 / 1M out |
| GPT-OSS 20B | OpenAI | 128K | $0.20 / 1M out |
| Kimi K2.6 | Moonshot AI | 256K | $4.00 / 1M out |
| Kimi K2.7 Code | Moonshot AI | 256K | $4.00 / 1M out |
| Kimi K3 | Moonshot AI | 1024K | $15.00 / 1M out |
| Llama 3.3 70B Instruct | Meta | 32K | $1.04 / 1M out |
| Minimax M2.7 | MiniMax | 192K | $1.20 / 1M out |
| Minimax M3 | MiniMax | 512K | $1.20 / 1M out |
| Muse Glimmer 30B | Meta | 128K | $1.50 / 1M out |
| Qwen 2.5 7B Instruct | Qwen | 8K | $0.30 / 1M out |
| Qwen3.5 9B | Qwen | 256K | $0.25 / 1M out |
| Qwen3.6 Plus | Qwen | 977K | $3.00 / 1M out |
| Qwen3.7 Max | Qwen | 977K | $3.75 / 1M out |
| Qwen3.7 Plus | Qwen | 977K | $1.28 / 1M out |
| Qwen3.8 2.4t A95B | Qwen | 986K | $6.25 / 1M out |
| Model | Author | Price |
|---|---|---|
| Flash Image 2.5 | $0.039 per image | |
| Flash Image 3.1 | $0.047 per image | |
| Flash Image 3.1 Lite | $0.069 per image | |
| FLUX.1 [schnell] | Black Forest Labs | $0.0027 per output megapixel |
| FLUX.1 Kontext Max | Black Forest Labs | $0.080 per output megapixel |
| FLUX.1 Kontext Pro | Black Forest Labs | $0.040 per output megapixel |
| FLUX.1.1 Pro | Black Forest Labs | $0.040 per output megapixel |
| FLUX.2 Dev | Black Forest Labs | $0.015 per image |
| FLUX.2 Flex | Black Forest Labs | $0.030 per image |
| FLUX.2 Max | Black Forest Labs | $0.070 per output megapixel |
| FLUX.2 Pro | Black Forest Labs | $0.030 per image |
| Gemini 3 Pro Image | $0.134 per image | |
| GPT Image 1.5 | OpenAI | $0.034 per image |
| GPT Image 2 | OpenAI | $0.053 per image |
| Imagen 4.0 Fast | $0.020 per image | |
| Imagen 4.0 Preview | $0.040 per image | |
| Imagen 4.0 Ultra | $0.060 per image | |
| Qwen Image | Qwen | $0.0058 per image |
| Qwen Image 2.0 | Qwen | $0.035 per image |
| Qwen Image 2.0 Pro | Qwen | $0.075 per image |
| Seedream 3.0 | ByteDance | $0.018 per image |
| Seedream 4.0 | ByteDance | $0.030 per image |
| Seedream 5.0 Lite | ByteDance | $0.035 per image |
| Wan2.6 Image | Wan | $0.030 per image |
| Model | Author | Price |
|---|---|---|
| FLUX 3 | Black Forest Labs | $0.170 per second of video |
| Hailuo 02 | MiniMax | $0.056 per second of video |
| Happyhorse 1.0 T2v | Alibaba | $0.240 per second of video |
| Happyhorse 1.1 T2v | Alibaba | $0.140 per second of video |
| Kling 1.6 Standard | Kling | $0.037 per second of video |
| Kling 2.1 Master | Kling | $0.185 per second of video |
| Kling 2.1 Pro | Kling | $0.065 per second of video |
| Kling 2.1 Standard | Kling | $0.037 per second of video |
| Pixverse V5 | PixVerse | $0.060 per second of video |
| Seedance 1.0 Lite | ByteDance | $0.029 per second of video |
| Seedance 1.0 Pro | ByteDance | $0.113 per second of video |
| Seedance 2.0 | ByteDance | $0.160 per second of video |
| Seedance 2.5 | ByteDance | $0.115 per second of video |
| Sora 2 | OpenAI | $0.100 per second of video |
| Sora 2 Pro | OpenAI | $0.375 per second of video |
| Veo 2.0 | $0.500 per second of video | |
| Veo 3.0 Audio | $0.400 per second of video | |
| Veo 3.0 Fast | $0.100 per second of video | |
| Veo 3.0 Fast Audio | $0.150 per second of video | |
| Vidu 2.0 | Vidu | $0.100 per second of video |
| Vidu Q1 | Vidu | $0.044 per second of video |
| Vidu Q3 | Vidu | $0.098 per second of video |
| Vidu Q3 Turbo | Vidu | $0.195 per second of video |
| Wan2.7 T2v | Wan | $0.020 per second of video |
Every model is served through one OpenAI-compatible endpoint at https://api.gpu.ai/v1. Chat models bill per input and output token; image and video models bill per generation. Prices are the live catalog rates — no subscriptions, no minimums, billed against your GPU.ai balance. Models marked economy trade a cold start of roughly 30–60 seconds for a lower rate. This page refreshes hourly from the live model catalog.