GPU.ai has a new look. The compute is the same.Read about the redesign  ↗

model apis · chat / llm

Muse Glimmer 30B API●

Run Muse Glimmer 30B by Meta through one OpenAI-compatible endpoint. Pay per use against your GPU.ai balance. No instance to rent, no deployment to manage.

Available now

Muse Glimmer 30B API pricing

Live catalog rates, billed per use against your GPU.ai balance. No subscriptions, no minimums.

TierRate
Serverless
Input tokens: $0.35 / 1MCached input tokens: $0.040 / 1MOutput tokens: $1.50 / 1M

rates are read every few minutes from the live model catalog and metered per request

Call Muse Glimmer 30B in one request

Any OpenAI SDK works: set the base URL to https://api.gpu.ai/v1 and use your GPU.ai API key.

curl
curl https://api.gpu.ai/v1/chat/completions \
  -H "Authorization: Bearer $GPUAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpuai/muse-glimmer-30b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Python (OpenAI SDK)
from openai import OpenAI

client = OpenAI(
    base_url="https://api.gpu.ai/v1",
    api_key="YOUR_GPUAI_API_KEY",
)

response = client.chat.completions.create(
    model="gpuai/muse-glimmer-30b",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
JavaScript (OpenAI SDK)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.gpu.ai/v1",
  apiKey: process.env.GPUAI_API_KEY,
});

const response = await client.chat.completions.create({
  model: "gpuai/muse-glimmer-30b",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);

Muse Glimmer 30B on GPU.ai

Model idgpuai/muse-glimmer-30b
ModalityChat / LLM
AuthorMeta
Context window131,072 tokens
Streaming—
Aliases—

faq

Muse Glimmer 30B API questions●

How much does the Muse Glimmer 30B API cost?

Muse Glimmer 30B on GPU.ai currently costs $0.35 / 1M (input tokens), $0.040 / 1M (cached input tokens), $1.50 / 1M (output tokens). There are no subscriptions or minimums. Usage is billed against your GPU.ai balance as you go.

Is the Muse Glimmer 30B API OpenAI-compatible?

Yes. Point any OpenAI SDK at https://api.gpu.ai/v1 with a GPU.ai API key and request model "gpuai/muse-glimmer-30b". No proprietary client needed.

What is the context window of Muse Glimmer 30B?

Muse Glimmer 30B supports a 131,072-token context window on GPU.ai.

How does billing work?

Chat models bill per input and output token at the listed per-million-token rates, metered per request against your GPU.ai balance. Where a cached-input rate is listed, the part of a prompt the serving partner answers from its prompt cache (reported as prompt_tokens_details.cached_tokens) bills at that lower rate instead of the input rate.