Now offering crypto payments →

Model APIs · Chat / LLM

Gemma 4 31B It API

Run Gemma 4 31B It by Google through one OpenAI-compatible endpoint. Pay per use against your GPU.ai balance — no instance to rent, no deployment to manage.

Available now
Pricing

Gemma 4 31B It API pricing

Live catalog rates, billed per use against your GPU.ai balance. No subscriptions, no minimums.

TierRate
Serverless
Input tokens: $0.39 / 1MOutput tokens: $0.97 / 1M

Rates are read hourly from the live model catalog and metered per request.

Quick start

Call Gemma 4 31B It in one request

Any OpenAI SDK works — set the base URL to https://api.gpu.ai/v1 and use your GPU.ai API key.

curl
curl https://api.gpu.ai/v1/chat/completions \
  -H "Authorization: Bearer $GPUAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpuai/gemma-4-31b-it",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Python (OpenAI SDK)
from openai import OpenAI

client = OpenAI(
    base_url="https://api.gpu.ai/v1",
    api_key="YOUR_GPUAI_API_KEY",
)

response = client.chat.completions.create(
    model="gpuai/gemma-4-31b-it",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
JavaScript (OpenAI SDK)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.gpu.ai/v1",
  apiKey: process.env.GPUAI_API_KEY,
});

const response = await client.chat.completions.create({
  model: "gpuai/gemma-4-31b-it",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);
Details

Gemma 4 31B It on GPU.ai

Model idgpuai/gemma-4-31b-it
ModalityChat / LLM
AuthorGoogle
Context window262,144 tokens
Streaming
Aliases
FAQ

Gemma 4 31B It API questions

How much does the Gemma 4 31B It API cost?

Gemma 4 31B It on GPU.ai currently costs $0.39 / 1M (input tokens), $0.97 / 1M (output tokens). There are no subscriptions or minimums — usage is billed against your GPU.ai balance as you go.

Is the Gemma 4 31B It API OpenAI-compatible?

Yes. Point any OpenAI SDK at https://api.gpu.ai/v1 with a GPU.ai API key and request model "gpuai/gemma-4-31b-it" — no proprietary client needed.

What is the context window of Gemma 4 31B It?

Gemma 4 31B It supports a 262,144-token context window on GPU.ai.

How does billing work?

Chat models bill per input and output token at the listed per-million-token rates, metered per request against your GPU.ai balance.