Qwen logo

Qwen3 Embedding 8B

QwenReleased Jun 4, 2025qwen/qwen3-embedding-8b

Qwen3-Embedding-8B is the largest model in the Qwen3 Embedding series, purpose-built for text embedding and retrieval tasks. It ranks No.1 on the MTEB multilingual leaderboard, supports 100+ languages, and produces embeddings up to 4096 dimensions with support for user-defined output dimensions (Matryoshka Representation Learning) from 32 to 4096. The model is instruction-aware, allowing custom task instructions to be prepended to inputs for improved retrieval performance, and supports a 32K token context length.

Context
33K
Max output
4K
Input /1M
$0.01
Output /1M
—
AcceptsTextProducesEmbeddingsInstruct type none

Providers

2 providers serve this model. FastRouter routes each request to the best available one, and you can bring your own key for any of them.

Fireworks AI logo
Fireworks AI
fireworks
Input /1M$0.10
Output /1M—
Context33K
Max output4K
DeepInfra logo
DeepInfra
deepinfra
Input /1M$0.01
Output /1M—
Context33K
Max output4K

Supported parameters

Request parameters you can send with this model.

Parameter
Type
Fireworks AI
DeepInfra
Other
dimensions
number
-
32 to 8192
normalize
bool
true | false
true | false
prompt_template
string
-
-
return_logits
array
-
encoding_format
string
-
"float", "base64"
inputs
array
-
instruction
string
-
-

Make your first API call

OpenAI-compatible. Point your SDK at FastRouter, or call the API directly.

Full API docs
Replace <FASTROUTER_API_KEY> with your key · Get your key →Install
OpenAI SDKHTTP
from openai import OpenAI
client = OpenAI(
base_url="https://api.fastrouter.ai/api/v1",
api_key="<FASTROUTER_API_KEY>",
)
embeddings = client.embeddings.create(
model="qwen/qwen3-embedding-8b",
input="What is node JS?",
dimensions=4
)
print("Embeddings:", embeddings)

Frequently asked questions

More models from Qwen

Qwen logo
Qwen2.5 72B InstructQwen/Qwen2.5-72B-Instruct

Alibaba's largest Qwen2.5 model featuring improved capabilities in coding, mathematics, and instruction following across more than 29 languages. With 72B parameters and 128K context support, it delivers top-tier performance for complex tasks requiring deep context understanding and sophisticated reasoning.

Qwen logo
Qwen3 14BQwen/Qwen3-14B

A versatile dense model offering powerful reasoning abilities paired with efficient dialogue processing. Developed by Alibaba, it supports extensive 32K context windows (extendable to 131K), features both thinking and non-thinking modes, and excels in coding, mathematics, and multilingual support across 119 languages and dialects.

Qwen logo
Qwen3 30B A3BQwen/Qwen3-30B-A3B

A Mixture-of-Experts model that activates only 3.3B of its 30.5B parameters per forward pass. Offers dual thinking modes: a detailed step-by-step reasoning mode for complex problems and a faster non-thinking mode for simpler queries. Despite its small active parameter count, it outperforms QwQ-32B that has 10 times more activated parameters.

Qwen logo
Qwen Imageqwen/qwen-image

Qwen-Image is an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. It supports parallel CFG and LoRA merging (up to 3 LoRAs), configurable inference steps, and an optional turbo mode for faster generation using optimized settings (10 steps, CFG=1.2).

Qwen logo
Qwen Image 2512qwen/qwen-image-2512

Qwen Image 2512 is an improved version of Qwen Image with better text rendering, finer natural textures, and more realistic human generation. This 20B MMDiT model achieved top ranking among open-source models after 10,000 blind comparison rounds on AI Arena, released December 31, 2025, and is licensed under Apache 2.0.

Qwen logo
Qwen Image 3.0qwen/qwen-image-3

Qwen-Image 3.0 is a general-purpose image generation and editing model that balances quality and speed. It excels at complex text rendering (multi-line, paragraph-level layouts), fine detail, and realistic textures, and supports both text-to-image (T2I) and image-to-image/editing (I2I) on DashScope's multimodal-generation endpoint.

Qwen logo
Qwen Image 3.0 Proqwen/qwen-image-3-pro

Qwen-Image 3.0 Pro is the high-quality tier of the Qwen-Image 3.0 series, offering stronger text rendering, more realistic textures, and better semantic adherence for both text-to-image (T2I) and image-to-image/editing (I2I). Output resolution ranges from 512x512 up to 2048x2048 (PNG).

Qwen logo
Qwen3 32Bqwen/qwen3-32b

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, coding, and logical inference, and a "non-thinking" mode for faster, general-purpose conversation. The model demonstrates strong performance in instruction-following, agent tool use, creative writing, and multilingual tasks across 100+ languages and dialects. It natively handles 32K token contexts and can extend to 131K tokens using YaRN-based scaling.

Model scores sourced from ArtificialAnalysis