Model Library/Qwen3 235B A22B Instruct 2507

Qwen3 235B A22B Instruct 2507

qwen/qwen3-235b-a22b-instruct-2507

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following, logical reasoning, math, code, and tool usage. The model supports a native 262K context length and does not implement "thinking mode" (<think> blocks). Compared to its base variant, this version delivers significant gains in knowledge coverage, long-context reasoning, coding benchmarks, and alignment with open-ended tasks. It is particularly strong on multilingual understanding, math reasoning (e.g., AIME, HMMT), and alignment evaluations like Arena-Hard and WritingBench.

Características

API serverless

Documentación

qwen/qwen3-235b-a22b-instruct-2507 is available via Novita's serverless API, where you pay per token. There are several ways to call the API, including OpenAI-compatible endpoints with exceptional reasoning performance.

Implementaciones bajo demanda

Documentación

On-demand deployments allow you to use qwen/qwen3-235b-a22b-instruct-2507 on dedicated GPUs with high-performance serving stack with high reliability and no rate limits.

Serverless disponible

Ejecuta consultas de inmediato, paga solo por el uso

Entrada$0.09 / M Tokens

Salida$0.58 / M Tokens

Usa los siguientes ejemplos de código para integrarte con nuestra API:

1from openai import OpenAI
2
3client = OpenAI(
4    api_key="<Your API Key>",
5    base_url="https://api.novita.ai/openai"
6)
7
8response = client.chat.completions.create(
9    model="qwen/qwen3-235b-a22b-instruct-2507",
10    messages=[
11        {"role": "system", "content": "You are a helpful assistant."},
12        {"role": "user", "content": "Hello, how are you?"}
13    ],
14    max_tokens=16384,
15    temperature=0.7
16)
17
18print(response.choices[0].message.content)

Información

Proveedor

Qwen

Cuantización

fp8

Funcionalidad compatible

Longitud del contexto

131072

Salida máxima

16384

Serverless

Compatible

Function Calling

Compatible

Structured Output

Compatible

Capacidades de entrada

text

Capacidades de salida

text

Todo lo que necesitas para crear IA de producción.

Más de 200 modelos, GPUs bajo demanda y entornos de ejecución seguros para agentes, unificados bajo una API. Gratis para empezar, escala a medida que creces.