We match your deposits, up to $150 in free credits. Ends October 1.Learn more  

model apis · embedding

Qwen3 Embedding 8B API

Run Qwen3 Embedding 8B by Qwen through one OpenAI-compatible endpoint. Pay per use against your GPU.ai balance. No instance to rent, no deployment to manage.

Available now

Qwen3 Embedding 8B API pricing

Live catalog rates, billed per use against your GPU.ai balance. No subscriptions, no minimums.

TierRate
Serverless
Input tokens: $0.10 / 1M

rates are read hourly from the live model catalog and metered per request

Call Qwen3 Embedding 8B in one request

Any OpenAI SDK works: set the base URL to https://api.gpu.ai/v1 and use your GPU.ai API key.

curl
curl https://api.gpu.ai/v1/chat/completions \
  -H "Authorization: Bearer $GPUAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpuai/qwen3-embedding-8b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Python (OpenAI SDK)
from openai import OpenAI

client = OpenAI(
    base_url="https://api.gpu.ai/v1",
    api_key="YOUR_GPUAI_API_KEY",
)

response = client.chat.completions.create(
    model="gpuai/qwen3-embedding-8b",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
JavaScript (OpenAI SDK)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.gpu.ai/v1",
  apiKey: process.env.GPUAI_API_KEY,
});

const response = await client.chat.completions.create({
  model: "gpuai/qwen3-embedding-8b",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);

Qwen3 Embedding 8B on GPU.ai

Model idgpuai/qwen3-embedding-8b
Modalityembedding
AuthorQwen
Context window40,960 tokens
Streaming
Aliases

faq

Qwen3 Embedding 8B API questions

How much does the Qwen3 Embedding 8B API cost?

Qwen3 Embedding 8B on GPU.ai currently costs $0.10 / 1M (input tokens). There are no subscriptions or minimums. Usage is billed against your GPU.ai balance as you go.

Is the Qwen3 Embedding 8B API OpenAI-compatible?

Yes. Point any OpenAI SDK at https://api.gpu.ai/v1 with a GPU.ai API key and request model "gpuai/qwen3-embedding-8b". No proprietary client needed.

How does billing work?

Generation models bill per output (per image, or per second of video) at the listed rate, metered per request against your GPU.ai balance.