Now offering crypto payments →

Model APIs · Chat / LLM

Qwen 2.5 7B Instruct API

Run Qwen 2.5 7B Instruct by Qwen through one OpenAI-compatible endpoint. Pay per use against your GPU.ai balance — no instance to rent, no deployment to manage.

Available now
Pricing

Qwen 2.5 7B Instruct API pricing

Live catalog rates, billed per use against your GPU.ai balance. No subscriptions, no minimums.

TierRate
Serverless
Input tokens: $0.30 / 1MOutput tokens: $0.30 / 1M

Rates are read hourly from the live model catalog and metered per request.

Quick start

Call Qwen 2.5 7B Instruct in one request

Any OpenAI SDK works — set the base URL to https://api.gpu.ai/v1 and use your GPU.ai API key.

curl
curl https://api.gpu.ai/v1/chat/completions \
  -H "Authorization: Bearer $GPUAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpuai/qwen2.5-7b-instruct",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Python (OpenAI SDK)
from openai import OpenAI

client = OpenAI(
    base_url="https://api.gpu.ai/v1",
    api_key="YOUR_GPUAI_API_KEY",
)

response = client.chat.completions.create(
    model="gpuai/qwen2.5-7b-instruct",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
JavaScript (OpenAI SDK)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.gpu.ai/v1",
  apiKey: process.env.GPUAI_API_KEY,
});

const response = await client.chat.completions.create({
  model: "gpuai/qwen2.5-7b-instruct",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);
Details

Qwen 2.5 7B Instruct on GPU.ai

Model idgpuai/qwen2.5-7b-instruct
ModalityChat / LLM
AuthorQwen
Context window8,192 tokens
StreamingSupported
Aliases
FAQ

Qwen 2.5 7B Instruct API questions

How much does the Qwen 2.5 7B Instruct API cost?

Qwen 2.5 7B Instruct on GPU.ai currently costs $0.30 / 1M (input tokens), $0.30 / 1M (output tokens). There are no subscriptions or minimums — usage is billed against your GPU.ai balance as you go.

Is the Qwen 2.5 7B Instruct API OpenAI-compatible?

Yes. Point any OpenAI SDK at https://api.gpu.ai/v1 with a GPU.ai API key and request model "gpuai/qwen2.5-7b-instruct" — no proprietary client needed.

What is the context window of Qwen 2.5 7B Instruct?

Qwen 2.5 7B Instruct supports a 8,192-token context window on GPU.ai.

How does billing work?

Chat models bill per input and output token at the listed per-million-token rates, metered per request against your GPU.ai balance.