Model APIs · Chat / LLM
Run Kimi K2.6 by Moonshot AI through one OpenAI-compatible endpoint. Pay per use against your GPU.ai balance — no instance to rent, no deployment to manage.
Live catalog rates, billed per use against your GPU.ai balance. No subscriptions, no minimums.
| Tier | Rate |
|---|---|
| Serverless | Input tokens: $0.95 / 1MOutput tokens: $4.00 / 1M |
Rates are read hourly from the live model catalog and metered per request.
Any OpenAI SDK works — set the base URL to https://api.gpu.ai/v1 and use your GPU.ai API key.
curl https://api.gpu.ai/v1/chat/completions \
-H "Authorization: Bearer $GPUAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpuai/kimi-k2.6",
"messages": [{"role": "user", "content": "Hello!"}]
}'from openai import OpenAI
client = OpenAI(
base_url="https://api.gpu.ai/v1",
api_key="YOUR_GPUAI_API_KEY",
)
response = client.chat.completions.create(
model="gpuai/kimi-k2.6",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.gpu.ai/v1",
apiKey: process.env.GPUAI_API_KEY,
});
const response = await client.chat.completions.create({
model: "gpuai/kimi-k2.6",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);| Model id | gpuai/kimi-k2.6 |
| Modality | Chat / LLM |
| Author | Moonshot AI |
| Context window | 262,144 tokens |
| Streaming | — |
| Aliases | — |
Kimi K2.6 on GPU.ai currently costs $0.95 / 1M (input tokens), $4.00 / 1M (output tokens). There are no subscriptions or minimums — usage is billed against your GPU.ai balance as you go.
Yes. Point any OpenAI SDK at https://api.gpu.ai/v1 with a GPU.ai API key and request model "gpuai/kimi-k2.6" — no proprietary client needed.
Kimi K2.6 supports a 262,144-token context window on GPU.ai.
Chat models bill per input and output token at the listed per-million-token rates, metered per request against your GPU.ai balance.