Now offering crypto payments →

Model APIs · Chat / LLM

GPT-OSS 120B API

Run GPT-OSS 120B by OpenAI through one OpenAI-compatible endpoint. Pay per use against your GPU.ai balance — no instance to rent, no deployment to manage.

Available now
Pricing

GPT-OSS 120B API pricing

Live catalog rates, billed per use against your GPU.ai balance. No subscriptions, no minimums.

TierRate
Serverless
Input tokens: $0.15 / 1MOutput tokens: $0.60 / 1M

Rates are read hourly from the live model catalog and metered per request.

Quick start

Call GPT-OSS 120B in one request

Any OpenAI SDK works — set the base URL to https://api.gpu.ai/v1 and use your GPU.ai API key.

curl
curl https://api.gpu.ai/v1/chat/completions \
  -H "Authorization: Bearer $GPUAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpuai/gpt-oss-120b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Python (OpenAI SDK)
from openai import OpenAI

client = OpenAI(
    base_url="https://api.gpu.ai/v1",
    api_key="YOUR_GPUAI_API_KEY",
)

response = client.chat.completions.create(
    model="gpuai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
JavaScript (OpenAI SDK)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.gpu.ai/v1",
  apiKey: process.env.GPUAI_API_KEY,
});

const response = await client.chat.completions.create({
  model: "gpuai/gpt-oss-120b",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);
Details

GPT-OSS 120B on GPU.ai

Model idgpuai/gpt-oss-120b
ModalityChat / LLM
AuthorOpenAI
Context window131,072 tokens
StreamingSupported
Aliases
FAQ

GPT-OSS 120B API questions

How much does the GPT-OSS 120B API cost?

GPT-OSS 120B on GPU.ai currently costs $0.15 / 1M (input tokens), $0.60 / 1M (output tokens). There are no subscriptions or minimums — usage is billed against your GPU.ai balance as you go.

Is the GPT-OSS 120B API OpenAI-compatible?

Yes. Point any OpenAI SDK at https://api.gpu.ai/v1 with a GPU.ai API key and request model "gpuai/gpt-oss-120b" — no proprietary client needed.

What is the context window of GPT-OSS 120B?

GPT-OSS 120B supports a 131,072-token context window on GPU.ai.

How does billing work?

Chat models bill per input and output token at the listed per-million-token rates, metered per request against your GPU.ai balance.