GPU.ai has a new look. The compute is the same.Read about the redesign  ↗

serverless inference

AI model API pricing●

Every model behind our OpenAI-compatible inference API, priced per use. Point your existing SDK at api.gpu.ai/v1 and pay per token or per image. No instance to rent, no cold infrastructure to babysit.

Chat / LLM models

13 models
ModelAuthorContextInput / 1MCached input / 1MOutput / 1M
GLM 5.2Z.AI198K$1.40$0.260$4.40
GLM 5.3Z.AI198K$1.40$0.260$4.40
GLM 5.3 FlashZ.AI1024K$0.15$0.030$0.50
GPT-OSS 120BOpenAI128K$0.15—$0.60
Kimi K3Moonshot AI1024K$2.70$0.270$13.50
Llama 3.3 70B InstructMeta32K$1.04—$1.04
Minimax M3MiniMax512K$0.30$0.060$1.20
Muse Glimmer 30BMeta128K$0.35$0.040$1.50
Qwen3.5 9BQwen256K$0.17—$0.25
Qwen3.6 PlusQwen977K$0.50—$3.00
Qwen3.7 PlusQwen977K$0.32—$1.28
Qwen3.8 2.4t A95BQwen986K$2.00$0.250$6.00
Qwen3.8 FlashQwen977K$0.15—$0.47

Image generation models

20 models
ModelAuthorPrice
Flash Image 2.5Google$0.039 per image
Flash Image 3.1Google$0.047 per image
Flash Image 3.1 LiteGoogle$0.069 per image
FLUX.1 Kontext MaxBlack Forest Labs$0.080 per output megapixel
FLUX.1 Kontext ProBlack Forest Labs$0.040 per output megapixel
FLUX.1.1 ProBlack Forest Labs$0.040 per output megapixel
FLUX.2 DevBlack Forest Labs$0.015 per image
FLUX.2 FlexBlack Forest Labs$0.030 per image
FLUX.2 MaxBlack Forest Labs$0.070 per output megapixel
FLUX.2 ProBlack Forest Labs$0.030 per image
Gemini 3 Pro ImageGoogle$0.134 per image
GPT Image 1.5OpenAI$0.034 per image
GPT Image 2OpenAI$0.053 per image
Qwen ImageQwen$0.0058 per image
Qwen Image 2.0Qwen$0.035 per image
Qwen Image 2.0 ProQwen$0.075 per image
Seedream 3.0ByteDance$0.018 per image
Seedream 4.0ByteDance$0.030 per image
Seedream 5.0 LiteByteDance$0.035 per image
Wan2.6 ImageWan$0.030 per image

Video generation models

14 models
ModelAuthorPrice
FLUX 3Black Forest Labs$0.170 per second of video
Hailuo 02MiniMax$0.056 per second of video
Happyhorse 1.0 T2vAlibaba$0.240 per second of video
Happyhorse 1.1 T2vAlibaba$0.140 per second of video
Minimax H3MiniMax$0.139 per second of video
Pixverse V5PixVerse$0.060 per second of video
Seedance 1.0 LiteByteDance$0.029 per second of video
Seedance 1.0 ProByteDance$0.113 per second of video
Seedance 2.0ByteDance$0.160 per second of video
Seedance 2.5ByteDance$0.115 per second of video
Vidu Q1Vidu$0.044 per second of video
Vidu Q3Vidu$0.098 per second of video
Vidu Q3 TurboVidu$0.195 per second of video
Wan2.7 T2vWan$0.020 per second of video

How this pricing works

Every model is served through one OpenAI-compatible endpoint at https://api.gpu.ai/v1. Chat models bill per input and output token; the part of a prompt a partner serves from its prompt cache bills at the lower cached-input rate where one is listed (a dash means the model has no cached rate and every prompt token bills at the input rate). Image and video models bill per generation. Prices are the live catalog rates, the same figures the API and the console bill against: no subscriptions, no minimums, billed against your GPU.ai balance. Models marked economy trade a cold start of roughly 30–60 seconds for a lower rate. This page refreshes every few minutes from the live model catalog.