Qwen3 32B

qwen/qwen3-32b-fp8

Achieves effective integration of inference and non-inference modes, allowing seamless switching between modes during conversations. Its inference capability matches that of QwQ-32B with a smaller parameter size, and its general capabilities significantly surpass those of Qwen2.5-14B, reaching the state-of-the-art (SOTA) level among models of the same scale.

Features

On-demand Deployments

Docs

On-demand deployments allow you to use qwen/qwen3-32b-fp8 on dedicated GPUs with high-performance serving stack with high reliability and no rate limits.

Info

Provider

Qwen

Quantization

fp8

Supported Functionality

Context Length

40960

Max Output

20000

Serverless

Not supported

Reasoning

Supported

Input Capabilities

text

Output Capabilities

text

Everything you need to build production AI.

200+ models, on-demand GPUs, and secure agent runtimes — unified under one API. Free to start, scales as you grow.