Model APIs · Chat / LLM
Run GPT-OSS 20B by OpenAI through one OpenAI-compatible endpoint. Pay per use against your GPU.ai balance — no instance to rent, no deployment to manage.
Live catalog rates, billed per use against your GPU.ai balance. No subscriptions, no minimums.
| Tier | Rate |
|---|---|
| Serverless | Input tokens: $0.0500 / 1MOutput tokens: $0.20 / 1M |
Rates are read hourly from the live model catalog and metered per request.
Any OpenAI SDK works — set the base URL to https://api.gpu.ai/v1 and use your GPU.ai API key.
curl https://api.gpu.ai/v1/chat/completions \
-H "Authorization: Bearer $GPUAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpuai/gpt-oss-20b",
"messages": [{"role": "user", "content": "Hello!"}]
}'from openai import OpenAI
client = OpenAI(
base_url="https://api.gpu.ai/v1",
api_key="YOUR_GPUAI_API_KEY",
)
response = client.chat.completions.create(
model="gpuai/gpt-oss-20b",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.gpu.ai/v1",
apiKey: process.env.GPUAI_API_KEY,
});
const response = await client.chat.completions.create({
model: "gpuai/gpt-oss-20b",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);| Model id | gpuai/gpt-oss-20b |
| Modality | Chat / LLM |
| Author | OpenAI |
| Context window | 131,072 tokens |
| Streaming | — |
| Aliases | — |
GPT-OSS 20B on GPU.ai currently costs $0.0500 / 1M (input tokens), $0.20 / 1M (output tokens). There are no subscriptions or minimums — usage is billed against your GPU.ai balance as you go.
Yes. Point any OpenAI SDK at https://api.gpu.ai/v1 with a GPU.ai API key and request model "gpuai/gpt-oss-20b" — no proprietary client needed.
GPT-OSS 20B supports a 131,072-token context window on GPU.ai.
Chat models bill per input and output token at the listed per-million-token rates, metered per request against your GPU.ai balance.