Model Library/MiMo V2.6 Flash
MiMo V2.6 Flash

MiMo V2.6 Flash

xiaomimimo/mimo-v2.6-flash
MiMo V2.6 Flash is Xiaomi's cost-efficient open-source reasoning model for high-frequency calls and large-scale professional workflows. It combines native text, image, video, and audio understanding with a 1M-token context window and up to 128K output tokens, while supporting controllable deep thinking, tool calling, JSON mode, streaming, and prompt caching.

Fonctionnalités

API sans serveur

Documentation

xiaomimimo/mimo-v2.6-flash is available via Novita's serverless API, where you pay per token. There are several ways to call the API, including OpenAI-compatible endpoints with exceptional reasoning performance.

Sans serveur disponible

Exécutez des requêtes immédiatement, ne payez que pour l’utilisation

Entrée$0.14 / M Tokens
Lecture du cache$0.0028 / M Tokens
Sortie$0.28 / M Tokens

Utilisez les exemples de code suivants pour intégrer notre API :

1from openai import OpenAI
2
3client = OpenAI(
4    api_key="<Your API Key>",
5    base_url="https://api.novita.ai/openai"
6)
7
8response = client.chat.completions.create(
9    model="xiaomimimo/mimo-v2.6-flash",
10    messages=[
11        {"role": "system", "content": "You are a helpful assistant."},
12        {"role": "user", "content": "Hello, how are you?"}
13    ],
14    max_tokens=131072,
15    temperature=0.7
16)
17
18print(response.choices[0].message.content)

Infos

Fournisseur
Xiaomi
Quantification
fp8

Fonctionnalités prises en charge

Longueur du contexte
1M
Sortie maximale
128K
Serverless
Pris en charge
Function Calling
Pris en charge
Structured Output
Pris en charge
Reasoning
Pris en charge
API Anthropic
Pris en charge
Capacités d’entrée
text, image, video, audio
Capacités de sortie
text

Tout ce dont vous avez besoin pour créer une IA de production.

Plus de 200 modèles, des GPUs à la demande et des environnements d’exécution d’agents sécurisés — unifiés sous une seule API. Gratuit pour commencer, évolutif à mesure que vous grandissez.