Model Library/Macaron V1 Venti
Mind Lab

Macaron V1 Venti

mindai/macaron-v1-venti
Macaron-V1-Venti is a 748B-parameter flagship model in the Macaron-V1 family, built for personal intelligence, tool use, coding workflows, and code-native Generative UI. The model uses a Mixture of LoRA (MoL) architecture on top of GLM-5.2, consisting of a 744B-parameter base model and four 1B-parameter LoRA specialists. The specialists cover chat, personal-agent tasks, coding, and GenUI, with an L0 router selecting the most suitable specialist for each new user request.

Features

Serverless API

Docs

mindai/macaron-v1-venti is available via Novita's serverless API, where you pay per token. There are several ways to call the API, including OpenAI-compatible endpoints with exceptional reasoning performance.

Available Serverless

Run queries immediately, pay only for usage

Input$0 / M Tokens
Output$0 / M Tokens

Use the following code examples to integrate with our API:

1from openai import OpenAI
2
3client = OpenAI(
4    api_key="<Your API Key>",
5    base_url="https://api.novita.ai/openai"
6)
7
8response = client.chat.completions.create(
9    model="mindai/macaron-v1-venti",
10    messages=[
11        {"role": "system", "content": "You are a helpful assistant."},
12        {"role": "user", "content": "Hello, how are you?"}
13    ],
14    max_tokens=131072,
15    temperature=0.7
16)
17
18print(response.choices[0].message.content)

Info

Provider
Mind Lab
Quantization
-

Supported Functionality

Context Length
1048576
Max Output
131072
Serverless
Supported
Function Calling
Supported
Reasoning
Supported
Anthropic API
Supported
Input Capabilities
text
Output Capabilities
text

Everything you need to build production AI.

200+ models, on-demand GPUs, and secure agent runtimes — unified under one API. Free to start, scales as you grow.