Model Library/Qwen3.8 Flash
Qwen

Qwen3.8 Flash

パートナー

qwen/qwen3.8-flash
Qwen3.8-Flash is Alibaba's multimodal MoE model and an early preview of the Qwen4 architecture: 125B total parameters with only 6B activated per token, plus a 51B N-gram embedding, built on GDN + QSA hybrid attention. It accepts text, image and video input across a 1M-token context and emits up to 131K tokens, with thinking mode on by default and switchable off. Built for coding, agentic workflows, visual and long-document understanding, and long-video analysis at a fraction of the cost of comparable frontier models.

機能

サーバーレス API

ドキュメント

qwen/qwen3.8-flash is available via Novita's serverless API, where you pay per token. There are several ways to call the API, including OpenAI-compatible endpoints with exceptional reasoning performance.

利用可能なサーバーレス

クエリをすぐに実行し、使用した分だけお支払い

入力$0.15 / M Tokens
キャッシュ読み取り$0.016 / M Tokens
出力$0.47 / M Tokens

以下のコード例を使用して、当社の API と統合してください:

1from openai import OpenAI
2
3client = OpenAI(
4    api_key="<Your API Key>",
5    base_url="https://api.novita.ai/openai"
6)
7
8response = client.chat.completions.create(
9    model="qwen/qwen3.8-flash",
10    messages=[
11        {"role": "system", "content": "You are a helpful assistant."},
12        {"role": "user", "content": "Hello, how are you?"}
13    ],
14    max_tokens=131072,
15    temperature=0.7
16)
17
18print(response.choices[0].message.content)

情報

プロバイダー
Qwen
量子化
-

サポートされている機能

コンテキスト長
977K
最大出力
128K
Serverless
サポートされています
Function Calling
サポートされています
Structured Output
サポートされています
Reasoning
サポートされています
Anthropic API
サポートされています
入力機能
text, image, video
出力機能
text

本番環境向けAIを構築するために必要なすべて。

200以上のモデル、オンデマンド GPUs、安全なエージェントランタイムを、1つの API に統合。無料で始められ、成長に合わせてスケールできます。