Model Library/MiniMax M2.5-highspeed

MiniMax M2.5-highspeed

minimax/minimax-m2.5-highspeed

MiniMax M2.5-highspeed is an accelerated SOTA model engineered for scenarios demanding extreme efficiency. It perfectly inherits the core intelligence and robust digital workspace capabilities of the standard M2.5—including its 80.2% score on SWE-Bench Verified, seamless manipulation of Office documents, and versatility in cross-software collaboration. With zero compromise on reasoning precision or logical depth, the Highspeed version delivers ultra-low latency inference through rigorous engineering optimization. This means you get more than just an intelligent assistant capable of planning and self-optimization; you gain a "high-velocity engine" that responds to high-frequency calls and processes complex document streams in near real-time, making it ideal for latency-sensitive interactive applications and large-scale automated pipelines.

機能

サーバーレス API

ドキュメント

minimax/minimax-m2.5-highspeed is available via Novita's serverless API, where you pay per token. There are several ways to call the API, including OpenAI-compatible endpoints with exceptional reasoning performance.

利用可能なサーバーレス

クエリをすぐに実行し、使用した分だけお支払い

入力$0.6 / M Tokens

キャッシュ読み取り$0.03 / M Tokens

出力$2.4 / M Tokens

以下のコード例を使用して、当社の API と統合してください:

1from openai import OpenAI
2
3client = OpenAI(
4    api_key="<Your API Key>",
5    base_url="https://api.novita.ai/openai"
6)
7
8response = client.chat.completions.create(
9    model="minimax/minimax-m2.5-highspeed",
10    messages=[
11        {"role": "system", "content": "You are a helpful assistant."},
12        {"role": "user", "content": "Hello, how are you?"}
13    ],
14    max_tokens=131100,
15    temperature=0.7
16)
17
18print(response.choices[0].message.content)

情報

プロバイダー

MiniMax

量子化

fp8

サポートされている機能

コンテキスト長

204800

最大出力

131100

Serverless

サポートされています

Function Calling

サポートされています

Structured Output

サポートされています

Reasoning

サポートされています

Anthropic API

サポートされています

入力機能

text

出力機能

text

本番環境向けAIを構築するために必要なすべて。

200以上のモデル、オンデマンド GPUs、安全なエージェントランタイムを、1つの API に統合。無料で始められ、成長に合わせてスケールできます。