# Pricing | Novita AI

> See pricing for 200+ AI models, GPU instances, and agent sandboxes. Developer-focused with startup-friendly rates. No hidden fees.

> For the complete documentation index, see [llms.txt](/llms.txt). Markdown is available with `Accept: text/markdown` and `.md` URL variants.

Source: /pricing

Model APIs

[Agent Sandbox](/sandbox)

GPUs

Resources

[Pricing](/pricing)

Start Building

![](/_next/image?url=%2Fpricing%2Fv5%2Fpricing-hero-bg.png&w=3840&q=75)

Pricing

# Pricing to seamlessly scale from idea to enterprise

Explore pricing for our Model APIs and GPU resources. Find the right plan to match your needs with transparent rates and flexible options.

Serverless EndpointsDedicated EndpointsAgent SandboxGPUs

Batch inference is available at an introductory 50% discount on input and output tokens for supported models. [Learn More](/docs/guides/llm-batch-api)

AllLLMImageAudioVideoAI SearchCache

All

![deepseek](/models/logo/svg/deepseek-logo.svg)Deepseek

Advanced AI models from DeepSeek, offering cutting-edge reasoning capabilities and competitive pricing for enterprise and research applications.

#### [DeepSeek V4 Pro 0813](/models/model-detail/deepseek-deepseek-v4-pro-0813?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v4-pro-0813?from=pricing)

Context1M

Input$1.32 /Mt· Cache Read $0.132 /Mt

Output$3.96 /Mt

#### [Deepseek V4 Flash 0731](/models/model-detail/deepseek-deepseek-v4-flash-0731?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v4-flash-0731?from=pricing)

Context1M

Input$0.44 /Mt· Cache Read $0.028 /Mt

Output$1.32 /Mt

#### [Deepseek V4 Flash](/models/model-detail/deepseek-deepseek-v4-flash?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v4-flash?from=pricing)

Context1M

Input$0.14 /Mt· Cache Read $0.028 /Mt

Output$0.28 /Mt

#### [Deepseek V4 Pro](/models/model-detail/deepseek-deepseek-v4-pro?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v4-pro?from=pricing)

Context1M

Input$1.6 /Mt· Cache Read $0.135 /Mt

Output$3.2 /Mt

#### [Deepseek V3.2](/models/model-detail/deepseek-deepseek-v3.2?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v3.2?from=pricing)

Context160K

Input$0.269 /Mt· Cache Read $0.1345 /Mt

Output$0.4 /Mt

#### [DeepSeek-OCR 2](/models/model-detail/deepseek-deepseek-ocr-2?from=pricing)

[More](/models/model-detail/deepseek-deepseek-ocr-2?from=pricing)

Context8K

Input$0.03 /Mt

Output$0.03 /Mt

#### [Deepseek V3.2 Exp](/models/model-detail/deepseek-deepseek-v3.2-exp?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v3.2-exp?from=pricing)

Context160K

Input$0.27 /Mt

Output$0.41 /Mt

#### [Deepseek V3.1 Terminus](/models/model-detail/deepseek-deepseek-v3.1-terminus?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v3.1-terminus?from=pricing)

Context128K

Input$0.27 /Mt· Cache Read $0.135 /Mt

Output$1 /Mt

#### [DeepSeek V3.1](/models/model-detail/deepseek-deepseek-v3.1?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v3.1?from=pricing)

Context128K

Input$0.27 /Mt· Cache Read $0.135 /Mt

Output$1 /Mt

#### [DeepSeek V3 0324](/models/model-detail/deepseek-deepseek-v3-0324?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v3-0324?from=pricing)

Context160K

Input$0.27 /Mt· Cache Read $0.135 /Mt

Output$1.12 /Mt

#### [DeepSeek R1 0528](/models/model-detail/deepseek-deepseek-r1-0528?from=pricing)

[More](/models/model-detail/deepseek-deepseek-r1-0528?from=pricing)

Context160K

Input$0.7 /Mt· Cache Read $0.35 /Mt

Output$2.5 /Mt

#### [DeepSeek R1 Distill LLama 70B](/models/model-detail/deepseek-deepseek-r1-distill-llama-70b?from=pricing)

[More](/models/model-detail/deepseek-deepseek-r1-distill-llama-70b?from=pricing)

Context8K

Input$0.8 /Mt

Output$0.8 /Mt

#### [DeepSeek V3 (Turbo)](/models/model-detail/deepseek-deepseek-v3-turbo?from=pricing)

[More](/models/model-detail/deepseek-deepseek-v3-turbo?from=pricing)

Context63K

Input$0.4 /Mt

Output$1.3 /Mt

#### [DeepSeek R1 (Turbo)](/models/model-detail/deepseek-deepseek-r1-turbo?from=pricing)

[More](/models/model-detail/deepseek-deepseek-r1-turbo?from=pricing)

Context63K

Input$0.7 /Mt

Output$2.5 /Mt

![deepseek](/models/logo/svg/deepseek-logo.svg)Deepseek

Advanced AI models from DeepSeek, offering cutting-edge reasoning capabilities and competitive pricing for enterprise and research applications.

Model NameContextInputOutputActions[DeepSeek V4 Pro 0813](/models/model-detail/deepseek-deepseek-v4-pro-0813?from=pricing)1M$1.32 /Mt· Cache Read $0.132 /Mt$3.96 /Mt[More](/models/model-detail/deepseek-deepseek-v4-pro-0813?from=pricing)[Deepseek V4 Flash 0731](/models/model-detail/deepseek-deepseek-v4-flash-0731?from=pricing)1M$0.44 /Mt· Cache Read $0.028 /Mt$1.32 /Mt[More](/models/model-detail/deepseek-deepseek-v4-flash-0731?from=pricing)[Deepseek V4 Flash](/models/model-detail/deepseek-deepseek-v4-flash?from=pricing)1M$0.14 /Mt· Cache Read $0.028 /Mt$0.28 /Mt[More](/models/model-detail/deepseek-deepseek-v4-flash?from=pricing)[Deepseek V4 Pro](/models/model-detail/deepseek-deepseek-v4-pro?from=pricing)1M$1.6 /Mt· Cache Read $0.135 /Mt$3.2 /Mt[More](/models/model-detail/deepseek-deepseek-v4-pro?from=pricing)[Deepseek V3.2](/models/model-detail/deepseek-deepseek-v3.2?from=pricing)160K$0.269 /Mt· Cache Read $0.1345 /Mt$0.4 /Mt[More](/models/model-detail/deepseek-deepseek-v3.2?from=pricing)[DeepSeek-OCR 2](/models/model-detail/deepseek-deepseek-ocr-2?from=pricing)8K$0.03 /Mt$0.03 /Mt[More](/models/model-detail/deepseek-deepseek-ocr-2?from=pricing)[Deepseek V3.2 Exp](/models/model-detail/deepseek-deepseek-v3.2-exp?from=pricing)160K$0.27 /Mt$0.41 /Mt[More](/models/model-detail/deepseek-deepseek-v3.2-exp?from=pricing)[Deepseek V3.1 Terminus](/models/model-detail/deepseek-deepseek-v3.1-terminus?from=pricing)128K$0.27 /Mt· Cache Read $0.135 /Mt$1 /Mt[More](/models/model-detail/deepseek-deepseek-v3.1-terminus?from=pricing)[DeepSeek V3.1](/models/model-detail/deepseek-deepseek-v3.1?from=pricing)128K$0.27 /Mt· Cache Read $0.135 /Mt$1 /Mt[More](/models/model-detail/deepseek-deepseek-v3.1?from=pricing)[DeepSeek V3 0324](/models/model-detail/deepseek-deepseek-v3-0324?from=pricing)160K$0.27 /Mt· Cache Read $0.135 /Mt$1.12 /Mt[More](/models/model-detail/deepseek-deepseek-v3-0324?from=pricing)[DeepSeek R1 0528](/models/model-detail/deepseek-deepseek-r1-0528?from=pricing)160K$0.7 /Mt· Cache Read $0.35 /Mt$2.5 /Mt[More](/models/model-detail/deepseek-deepseek-r1-0528?from=pricing)[DeepSeek R1 Distill LLama 70B](/models/model-detail/deepseek-deepseek-r1-distill-llama-70b?from=pricing)8K$0.8 /Mt$0.8 /Mt[More](/models/model-detail/deepseek-deepseek-r1-distill-llama-70b?from=pricing)[DeepSeek V3 (Turbo)](/models/model-detail/deepseek-deepseek-v3-turbo?from=pricing)63K$0.4 /Mt$1.3 /Mt[More](/models/model-detail/deepseek-deepseek-v3-turbo?from=pricing)[DeepSeek R1 (Turbo)](/models/model-detail/deepseek-deepseek-r1-turbo?from=pricing)63K$0.7 /Mt$2.5 /Mt[More](/models/model-detail/deepseek-deepseek-r1-turbo?from=pricing)

![qwen](/models/logo/svg/qwen-logo.svg)Qwen

Qwen series models offering efficient language processing with various parameter sizes, from lightweight to enterprise-grade solutions.

#### [Qwen3.8 Max](/models/model-detail/qwen-qwen3.8-max?from=pricing)

[More](/models/model-detail/qwen-qwen3.8-max?from=pricing)

Context977K

Input$2 /Mt· Cache Read $0.25 /Mt

Output$6 /Mt

#### [Qwen3.7-Max](/models/model-detail/qwen-qwen3.7-max?from=pricing)

[More](/models/model-detail/qwen-qwen3.7-max?from=pricing)

Context977K

Input$1.25 /Mt· Cache Read $0.25 /Mt

Output$3.75 /Mt

#### [Qwen3.6-27B](/models/model-detail/qwen-qwen3.6-27b?from=pricing)

[More](/models/model-detail/qwen-qwen3.6-27b?from=pricing)

Context256K

Input$0.6 /Mt

Output$3.6 /Mt

#### [Qwen3.5-27B](/models/model-detail/qwen-qwen3.5-27b?from=pricing)

[More](/models/model-detail/qwen-qwen3.5-27b?from=pricing)

Context256K

Input$0.3 /Mt

Output$2.4 /Mt

#### [Qwen3.5-122B-A10B](/models/model-detail/qwen-qwen3.5-122b-a10b?from=pricing)

[More](/models/model-detail/qwen-qwen3.5-122b-a10b?from=pricing)

Context256K

Input$0.4 /Mt

Output$3.2 /Mt

#### [Qwen3.5-35B-A3B](/models/model-detail/qwen-qwen3.5-35b-a3b?from=pricing)

[More](/models/model-detail/qwen-qwen3.5-35b-a3b?from=pricing)

Context256K

Input$0.25 /Mt

Output$2 /Mt

#### [Qwen3.5-397B-A17B](/models/model-detail/qwen-qwen3.5-397b-a17b?from=pricing)

[More](/models/model-detail/qwen-qwen3.5-397b-a17b?from=pricing)

Context256K

Input$0.6 /Mt

Output$3.6 /Mt

#### [Qwen3 Coder Next](/models/model-detail/qwen-qwen3-coder-next?from=pricing)

[More](/models/model-detail/qwen-qwen3-coder-next?from=pricing)

Context256K

Input$0.2 /Mt

Output$1.5 /Mt

#### [Qwen3 VL 235B A22B Thinking](/models/model-detail/qwen-qwen3-vl-235b-a22b-thinking?from=pricing)

[More](/models/model-detail/qwen-qwen3-vl-235b-a22b-thinking?from=pricing)

Context128K

Input$0.98 /Mt

Output$3.95 /Mt

#### [Qwen3.6-35B-A3B](/models/model-detail/qwen-qwen3.6-35b-a3b?from=pricing)

[More](/models/model-detail/qwen-qwen3.6-35b-a3b?from=pricing)

Context256K

Input$0.248 /Mt

Output$1.485 /Mt

#### [Qwen3 Next 80B A3B Instruct](/models/model-detail/qwen-qwen3-next-80b-a3b-instruct?from=pricing)

[More](/models/model-detail/qwen-qwen3-next-80b-a3b-instruct?from=pricing)

Context128K

Input$0.15 /Mt

Output$1.5 /Mt

#### [Qwen3 VL 235B A22B Instruct](/models/model-detail/qwen-qwen3-vl-235b-a22b-instruct?from=pricing)

[More](/models/model-detail/qwen-qwen3-vl-235b-a22b-instruct?from=pricing)

Context128K

Input$0.3 /Mt

Output$1.5 /Mt

#### [Qwen3 Max](/models/model-detail/qwen-qwen3-max?from=pricing)

[More](/models/model-detail/qwen-qwen3-max?from=pricing)

Context256K

Input-

OutputTiered pricing

#### [Qwen3 Coder 480B A35B Instruct](/models/model-detail/qwen-qwen3-coder-480b-a35b-instruct?from=pricing)

[More](/models/model-detail/qwen-qwen3-coder-480b-a35b-instruct?from=pricing)

Context256K

Input$0.38 /Mt

Output$1.55 /Mt

#### [Qwen3 Coder 30b A3B Instruct](/models/model-detail/qwen-qwen3-coder-30b-a3b-instruct?from=pricing)

[More](/models/model-detail/qwen-qwen3-coder-30b-a3b-instruct?from=pricing)

Context156K

Input$0.07 /Mt

Output$0.27 /Mt

#### [Qwen3 235B A22b Thinking 2507](/models/model-detail/qwen-qwen3-235b-a22b-thinking-2507?from=pricing)

[More](/models/model-detail/qwen-qwen3-235b-a22b-thinking-2507?from=pricing)

Context128K

Input$0.3 /Mt

Output$3 /Mt

#### [Qwen3 235B A22B Instruct 2507](/models/model-detail/qwen-qwen3-235b-a22b-instruct-2507?from=pricing)

[More](/models/model-detail/qwen-qwen3-235b-a22b-instruct-2507?from=pricing)

Context128K

Input$0.09 /Mt

Output$0.58 /Mt

#### [Qwen 2.5 72B Instruct](/models/model-detail/qwen-qwen-2.5-72b-instruct?from=pricing)

[More](/models/model-detail/qwen-qwen-2.5-72b-instruct?from=pricing)

Context31K

Input$0.38 /Mt

Output$0.4 /Mt

#### [Qwen3 235B A22B](/models/model-detail/qwen-qwen3-235b-a22b-fp8?from=pricing)

[More](/models/model-detail/qwen-qwen3-235b-a22b-fp8?from=pricing)

Context40K

Input$0.2 /Mt

Output$0.8 /Mt

#### [qwen/qwen3-vl-30b-a3b-instruct](/models/model-detail/qwen-qwen3-vl-30b-a3b-instruct?from=pricing)

[More](/models/model-detail/qwen-qwen3-vl-30b-a3b-instruct?from=pricing)

Context128K

Input$0.2 /Mt

Output$0.7 /Mt

#### [Qwen3 Omni 30B A3B Thinking](/models/model-detail/qwen-qwen3-omni-30b-a3b-thinking?from=pricing)

[More](/models/model-detail/qwen-qwen3-omni-30b-a3b-thinking?from=pricing)

Context64K

Input-

OutputOmnimodal

#### [Qwen3 Omni 30B A3B Instruct](/models/model-detail/qwen-qwen3-omni-30b-a3b-instruct?from=pricing)

[More](/models/model-detail/qwen-qwen3-omni-30b-a3b-instruct?from=pricing)

Context64K

Input-

OutputOmnimodal

#### [Qwen MT Plus](/models/model-detail/qwen-qwen-mt-plus?from=pricing)

[More](/models/model-detail/qwen-qwen-mt-plus?from=pricing)

Context16K

Input$0.25 /Mt

Output$0.75 /Mt

![qwen](/models/logo/svg/qwen-logo.svg)Qwen

Qwen series models offering efficient language processing with various parameter sizes, from lightweight to enterprise-grade solutions.

Model NameContextInputOutputActions[Qwen3.8 Max](/models/model-detail/qwen-qwen3.8-max?from=pricing)977K$2 /Mt· Cache Read $0.25 /Mt$6 /Mt[More](/models/model-detail/qwen-qwen3.8-max?from=pricing)[Qwen3.7-Max](/models/model-detail/qwen-qwen3.7-max?from=pricing)977K$1.25 /Mt· Cache Read $0.25 /Mt$3.75 /Mt[More](/models/model-detail/qwen-qwen3.7-max?from=pricing)[Qwen3.6-27B](/models/model-detail/qwen-qwen3.6-27b?from=pricing)256K$0.6 /Mt$3.6 /Mt[More](/models/model-detail/qwen-qwen3.6-27b?from=pricing)[Qwen3.5-27B](/models/model-detail/qwen-qwen3.5-27b?from=pricing)256K$0.3 /Mt$2.4 /Mt[More](/models/model-detail/qwen-qwen3.5-27b?from=pricing)[Qwen3.5-122B-A10B](/models/model-detail/qwen-qwen3.5-122b-a10b?from=pricing)256K$0.4 /Mt$3.2 /Mt[More](/models/model-detail/qwen-qwen3.5-122b-a10b?from=pricing)[Qwen3.5-35B-A3B](/models/model-detail/qwen-qwen3.5-35b-a3b?from=pricing)256K$0.25 /Mt$2 /Mt[More](/models/model-detail/qwen-qwen3.5-35b-a3b?from=pricing)[Qwen3.5-397B-A17B](/models/model-detail/qwen-qwen3.5-397b-a17b?from=pricing)256K$0.6 /Mt$3.6 /Mt[More](/models/model-detail/qwen-qwen3.5-397b-a17b?from=pricing)[Qwen3 Coder Next](/models/model-detail/qwen-qwen3-coder-next?from=pricing)256K$0.2 /Mt$1.5 /Mt[More](/models/model-detail/qwen-qwen3-coder-next?from=pricing)[Qwen3 VL 235B A22B Thinking](/models/model-detail/qwen-qwen3-vl-235b-a22b-thinking?from=pricing)128K$0.98 /Mt$3.95 /Mt[More](/models/model-detail/qwen-qwen3-vl-235b-a22b-thinking?from=pricing)[Qwen3.6-35B-A3B](/models/model-detail/qwen-qwen3.6-35b-a3b?from=pricing)256K$0.248 /Mt$1.485 /Mt[More](/models/model-detail/qwen-qwen3.6-35b-a3b?from=pricing)[Qwen3 Next 80B A3B Instruct](/models/model-detail/qwen-qwen3-next-80b-a3b-instruct?from=pricing)128K$0.15 /Mt$1.5 /Mt[More](/models/model-detail/qwen-qwen3-next-80b-a3b-instruct?from=pricing)[Qwen3 VL 235B A22B Instruct](/models/model-detail/qwen-qwen3-vl-235b-a22b-instruct?from=pricing)128K$0.3 /Mt$1.5 /Mt[More](/models/model-detail/qwen-qwen3-vl-235b-a22b-instruct?from=pricing)[Qwen3 Max](/models/model-detail/qwen-qwen3-max?from=pricing)256K-Tiered pricing[More](/models/model-detail/qwen-qwen3-max?from=pricing)[Qwen3 Coder 480B A35B Instruct](/models/model-detail/qwen-qwen3-coder-480b-a35b-instruct?from=pricing)256K$0.38 /Mt$1.55 /Mt[More](/models/model-detail/qwen-qwen3-coder-480b-a35b-instruct?from=pricing)[Qwen3 Coder 30b A3B Instruct](/models/model-detail/qwen-qwen3-coder-30b-a3b-instruct?from=pricing)156K$0.07 /Mt$0.27 /Mt[More](/models/model-detail/qwen-qwen3-coder-30b-a3b-instruct?from=pricing)[Qwen3 235B A22b Thinking 2507](/models/model-detail/qwen-qwen3-235b-a22b-thinking-2507?from=pricing)128K$0.3 /Mt$3 /Mt[More](/models/model-detail/qwen-qwen3-235b-a22b-thinking-2507?from=pricing)[Qwen3 235B A22B Instruct 2507](/models/model-detail/qwen-qwen3-235b-a22b-instruct-2507?from=pricing)128K$0.09 /Mt$0.58 /Mt[More](/models/model-detail/qwen-qwen3-235b-a22b-instruct-2507?from=pricing)[Qwen 2.5 72B Instruct](/models/model-detail/qwen-qwen-2.5-72b-instruct?from=pricing)31K$0.38 /Mt$0.4 /Mt[More](/models/model-detail/qwen-qwen-2.5-72b-instruct?from=pricing)[Qwen3 235B A22B](/models/model-detail/qwen-qwen3-235b-a22b-fp8?from=pricing)40K$0.2 /Mt$0.8 /Mt[More](/models/model-detail/qwen-qwen3-235b-a22b-fp8?from=pricing)[qwen/qwen3-vl-30b-a3b-instruct](/models/model-detail/qwen-qwen3-vl-30b-a3b-instruct?from=pricing)128K$0.2 /Mt$0.7 /Mt[More](/models/model-detail/qwen-qwen3-vl-30b-a3b-instruct?from=pricing)[Qwen3 Omni 30B A3B Thinking](/models/model-detail/qwen-qwen3-omni-30b-a3b-thinking?from=pricing)64K-Omnimodal[More](/models/model-detail/qwen-qwen3-omni-30b-a3b-thinking?from=pricing)[Qwen3 Omni 30B A3B Instruct](/models/model-detail/qwen-qwen3-omni-30b-a3b-instruct?from=pricing)64K-Omnimodal[More](/models/model-detail/qwen-qwen3-omni-30b-a3b-instruct?from=pricing)[Qwen MT Plus](/models/model-detail/qwen-qwen-mt-plus?from=pricing)16K$0.25 /Mt$0.75 /Mt[More](/models/model-detail/qwen-qwen-mt-plus?from=pricing)

Baidu

Baidu's ERNIE models providing advanced Chinese language understanding and multimodal capabilities, optimized for Chinese applications with competitive pricing.

#### [CoBuddy](/models/model-detail/baidu-cobuddy?from=pricing)

[More](/models/model-detail/baidu-cobuddy?from=pricing)

Context128K

Input$0.28 /Mt· Cache Read $0.07 /Mt

Output$1.13 /Mt

#### [ERNIE 4.5 VL 424B A47B](/models/model-detail/baidu-ernie-4.5-vl-424b-a47b?from=pricing)

[More](/models/model-detail/baidu-ernie-4.5-vl-424b-a47b?from=pricing)

Context120K

Input$0.42 /Mt

Output$1.25 /Mt

#### [ERNIE 4.5 21B A3B](/models/model-detail/baidu-ernie-4.5-21B-a3b?from=pricing)

[More](/models/model-detail/baidu-ernie-4.5-21B-a3b?from=pricing)

Context117K

Input$0.07 /Mt

Output$0.28 /Mt

Baidu

Baidu's ERNIE models providing advanced Chinese language understanding and multimodal capabilities, optimized for Chinese applications with competitive pricing.

Model NameContextInputOutputActions[CoBuddy](/models/model-detail/baidu-cobuddy?from=pricing)128K$0.28 /Mt· Cache Read $0.07 /Mt$1.13 /Mt[More](/models/model-detail/baidu-cobuddy?from=pricing)[ERNIE 4.5 VL 424B A47B](/models/model-detail/baidu-ernie-4.5-vl-424b-a47b?from=pricing)120K$0.42 /Mt$1.25 /Mt[More](/models/model-detail/baidu-ernie-4.5-vl-424b-a47b?from=pricing)[ERNIE 4.5 21B A3B](/models/model-detail/baidu-ernie-4.5-21B-a3b?from=pricing)117K$0.07 /Mt$0.28 /Mt[More](/models/model-detail/baidu-ernie-4.5-21B-a3b?from=pricing)

![zai-org](/models/logo/svg/glm-logo.svg)Zai-org

GLM series models from Tsinghua University, featuring advanced Chinese language understanding and generation capabilities.

#### [GLM 5.2](/models/model-detail/zai-org-glm-5.2?from=pricing)

[More](/models/model-detail/zai-org-glm-5.2?from=pricing)

Context1M

Input$1.4 /Mt· Cache Read $0.26 /Mt

Output$4.4 /Mt

#### [GLM-5.1](/models/model-detail/zai-org-glm-5.1?from=pricing)

[More](/models/model-detail/zai-org-glm-5.1?from=pricing)

Context200K

Input$1.38 /Mt· Cache Read $0.26 /Mt

Output$4.4 /Mt

#### [GLM-5](/models/model-detail/zai-org-glm-5?from=pricing)

[More](/models/model-detail/zai-org-glm-5?from=pricing)

Context198K

Input$1 /Mt· Cache Read $0.2 /Mt

Output$3.2 /Mt

#### [GLM-4.7-Flash](/models/model-detail/zai-org-glm-4.7-flash?from=pricing)

[More](/models/model-detail/zai-org-glm-4.7-flash?from=pricing)

Context195K

Input$0.07 /Mt· Cache Read $0.01 /Mt

Output$0.4 /Mt

#### [GLM-4.7](/models/model-detail/zai-org-glm-4.7?from=pricing)

[More](/models/model-detail/zai-org-glm-4.7?from=pricing)

Context200K

Input$0.6 /Mt· Cache Read $0.11 /Mt

Output$2.2 /Mt

#### [AutoGLM-Phone-9B-Multilingual](/models/model-detail/zai-org-autoglm-phone-9b-multilingual?from=pricing)

[More](/models/model-detail/zai-org-autoglm-phone-9b-multilingual?from=pricing)

Context64K

Input$0.035 /Mt

Output$0.138 /Mt

#### [GLM 4.6V](/models/model-detail/zai-org-glm-4.6v?from=pricing)

[More](/models/model-detail/zai-org-glm-4.6v?from=pricing)

Context128K

Input$0.3 /Mt· Cache Read $0.055 /Mt

Output$0.9 /Mt

#### [GLM 4.6](/models/model-detail/zai-org-glm-4.6?from=pricing)

[More](/models/model-detail/zai-org-glm-4.6?from=pricing)

Context200K

Input$0.55 /Mt· Cache Read $0.11 /Mt

Output$2.2 /Mt

#### [GLM 4.5V](/models/model-detail/zai-org-glm-4.5v?from=pricing)

[More](/models/model-detail/zai-org-glm-4.5v?from=pricing)

Context64K

Input$0.6 /Mt· Cache Read $0.11 /Mt

Output$1.8 /Mt

#### [zai-org/glm-4.5-air](/models/model-detail/zai-org-glm-4.5-air?from=pricing)

[More](/models/model-detail/zai-org-glm-4.5-air?from=pricing)

Context128K

Input$0.13 /Mt· Cache Read $0.025 /Mt

Output$0.85 /Mt

![zai-org](/models/logo/svg/glm-logo.svg)Zai-org

GLM series models from Tsinghua University, featuring advanced Chinese language understanding and generation capabilities.

Model NameContextInputOutputActions[GLM 5.2](/models/model-detail/zai-org-glm-5.2?from=pricing)1M$1.4 /Mt· Cache Read $0.26 /Mt$4.4 /Mt[More](/models/model-detail/zai-org-glm-5.2?from=pricing)[GLM-5.1](/models/model-detail/zai-org-glm-5.1?from=pricing)200K$1.38 /Mt· Cache Read $0.26 /Mt$4.4 /Mt[More](/models/model-detail/zai-org-glm-5.1?from=pricing)[GLM-5](/models/model-detail/zai-org-glm-5?from=pricing)198K$1 /Mt· Cache Read $0.2 /Mt$3.2 /Mt[More](/models/model-detail/zai-org-glm-5?from=pricing)[GLM-4.7-Flash](/models/model-detail/zai-org-glm-4.7-flash?from=pricing)195K$0.07 /Mt· Cache Read $0.01 /Mt$0.4 /Mt[More](/models/model-detail/zai-org-glm-4.7-flash?from=pricing)[GLM-4.7](/models/model-detail/zai-org-glm-4.7?from=pricing)200K$0.6 /Mt· Cache Read $0.11 /Mt$2.2 /Mt[More](/models/model-detail/zai-org-glm-4.7?from=pricing)[AutoGLM-Phone-9B-Multilingual](/models/model-detail/zai-org-autoglm-phone-9b-multilingual?from=pricing)64K$0.035 /Mt$0.138 /Mt[More](/models/model-detail/zai-org-autoglm-phone-9b-multilingual?from=pricing)[GLM 4.6V](/models/model-detail/zai-org-glm-4.6v?from=pricing)128K$0.3 /Mt· Cache Read $0.055 /Mt$0.9 /Mt[More](/models/model-detail/zai-org-glm-4.6v?from=pricing)[GLM 4.6](/models/model-detail/zai-org-glm-4.6?from=pricing)200K$0.55 /Mt· Cache Read $0.11 /Mt$2.2 /Mt[More](/models/model-detail/zai-org-glm-4.6?from=pricing)[GLM 4.5V](/models/model-detail/zai-org-glm-4.5v?from=pricing)64K$0.6 /Mt· Cache Read $0.11 /Mt$1.8 /Mt[More](/models/model-detail/zai-org-glm-4.5v?from=pricing)[zai-org/glm-4.5-air](/models/model-detail/zai-org-glm-4.5-air?from=pricing)128K$0.13 /Mt· Cache Read $0.025 /Mt$0.85 /Mt[More](/models/model-detail/zai-org-glm-4.5-air?from=pricing)

![sao10k](/_next/image?url=%2Fmodels%2Flogo%2Fsao10k-logo.png&w=48&q=75)Sao10K

Specialized fine-tuned models optimized for creative and roleplay applications with enhanced storytelling capabilities.

#### [Sao10k L3 8B Lunaris](/models/model-detail/sao10k-l3-8b-lunaris?from=pricing)

[More](/models/model-detail/sao10k-l3-8b-lunaris?from=pricing)

Context8K

Input$0.05 /Mt

Output$0.05 /Mt

#### [L3 8B Stheno V3.2](/models/model-detail/Sao10K-L3-8B-Stheno-v3.2?from=pricing)

[More](/models/model-detail/Sao10K-L3-8B-Stheno-v3.2?from=pricing)

Context8K

Input$0.05 /Mt

Output$0.05 /Mt

#### [L31 70B Euryale V2.2](/models/model-detail/sao10k-l31-70b-euryale-v2.2?from=pricing)

[More](/models/model-detail/sao10k-l31-70b-euryale-v2.2?from=pricing)

Context8K

Input$1.48 /Mt

Output$1.48 /Mt

![sao10k](/_next/image?url=%2Fmodels%2Flogo%2Fsao10k-logo.png&w=48&q=75)Sao10K

Specialized fine-tuned models optimized for creative and roleplay applications with enhanced storytelling capabilities.

Model NameContextInputOutputActions[Sao10k L3 8B Lunaris](/models/model-detail/sao10k-l3-8b-lunaris?from=pricing)8K$0.05 /Mt$0.05 /Mt[More](/models/model-detail/sao10k-l3-8b-lunaris?from=pricing)[L3 8B Stheno V3.2](/models/model-detail/Sao10K-L3-8B-Stheno-v3.2?from=pricing)8K$0.05 /Mt$0.05 /Mt[More](/models/model-detail/Sao10K-L3-8B-Stheno-v3.2?from=pricing)[L31 70B Euryale V2.2](/models/model-detail/sao10k-l31-70b-euryale-v2.2?from=pricing)8K$1.48 /Mt$1.48 /Mt[More](/models/model-detail/sao10k-l31-70b-euryale-v2.2?from=pricing)

MoonshotAI

--

#### [Kimi K3](/models/model-detail/moonshotai-kimi-k3?from=pricing)

[More](/models/model-detail/moonshotai-kimi-k3?from=pricing)

Context1M

Input$3 /Mt· Cache Read $0.3 /Mt

Output$15 /Mt

#### [Kimi K2.7 Code](/models/model-detail/moonshotai-kimi-k2.7-code?from=pricing)

[More](/models/model-detail/moonshotai-kimi-k2.7-code?from=pricing)

Context256K

Input$0.95 /Mt· Cache Read $0.19 /Mt

Output$4 /Mt

#### [Kimi K2.6](/models/model-detail/moonshotai-kimi-k2.6?from=pricing)

[More](/models/model-detail/moonshotai-kimi-k2.6?from=pricing)

Context256K

Input$0.8 /Mt· Cache Read $0.16 /Mt

Output$3.4 /Mt

#### [Kimi K2.5](/models/model-detail/moonshotai-kimi-k2.5?from=pricing)

[More](/models/model-detail/moonshotai-kimi-k2.5?from=pricing)

Context256K

Input$0.6 /Mt· Cache Read $0.1 /Mt

Output$3 /Mt

#### [Kimi K2 Thinking](/models/model-detail/moonshotai-kimi-k2-thinking?from=pricing)

[More](/models/model-detail/moonshotai-kimi-k2-thinking?from=pricing)

Context256K

Input$0.6 /Mt· Cache Read $0.15 /Mt

Output$2.5 /Mt

#### [Kimi K2 0905](/models/model-detail/moonshotai-kimi-k2-0905?from=pricing)

[More](/models/model-detail/moonshotai-kimi-k2-0905?from=pricing)

Context256K

Input$0.6 /Mt

Output$2.5 /Mt

#### [Kimi K2 Instruct](/models/model-detail/moonshotai-kimi-k2-instruct?from=pricing)

[More](/models/model-detail/moonshotai-kimi-k2-instruct?from=pricing)

Context128K

Input$0.57 /Mt

Output$2.3 /Mt

MoonshotAI

--

Model NameContextInputOutputActions[Kimi K3](/models/model-detail/moonshotai-kimi-k3?from=pricing)1M$3 /Mt· Cache Read $0.3 /Mt$15 /Mt[More](/models/model-detail/moonshotai-kimi-k3?from=pricing)[Kimi K2.7 Code](/models/model-detail/moonshotai-kimi-k2.7-code?from=pricing)256K$0.95 /Mt· Cache Read $0.19 /Mt$4 /Mt[More](/models/model-detail/moonshotai-kimi-k2.7-code?from=pricing)[Kimi K2.6](/models/model-detail/moonshotai-kimi-k2.6?from=pricing)256K$0.8 /Mt· Cache Read $0.16 /Mt$3.4 /Mt[More](/models/model-detail/moonshotai-kimi-k2.6?from=pricing)[Kimi K2.5](/models/model-detail/moonshotai-kimi-k2.5?from=pricing)256K$0.6 /Mt· Cache Read $0.1 /Mt$3 /Mt[More](/models/model-detail/moonshotai-kimi-k2.5?from=pricing)[Kimi K2 Thinking](/models/model-detail/moonshotai-kimi-k2-thinking?from=pricing)256K$0.6 /Mt· Cache Read $0.15 /Mt$2.5 /Mt[More](/models/model-detail/moonshotai-kimi-k2-thinking?from=pricing)[Kimi K2 0905](/models/model-detail/moonshotai-kimi-k2-0905?from=pricing)256K$0.6 /Mt$2.5 /Mt[More](/models/model-detail/moonshotai-kimi-k2-0905?from=pricing)[Kimi K2 Instruct](/models/model-detail/moonshotai-kimi-k2-instruct?from=pricing)128K$0.57 /Mt$2.3 /Mt[More](/models/model-detail/moonshotai-kimi-k2-instruct?from=pricing)

Hunyuan

--

#### [Hy3](/models/model-detail/tencent-hy3?from=pricing)

[More](/models/model-detail/tencent-hy3?from=pricing)

Context256K

Input$0.14 /Mt· Cache Read $0.035 /Mt

Output$0.58 /Mt

Hunyuan

--

Model NameContextInputOutputActions[Hy3](/models/model-detail/tencent-hy3?from=pricing)256K$0.14 /Mt· Cache Read $0.035 /Mt$0.58 /Mt[More](/models/model-detail/tencent-hy3?from=pricing)

![mind lab](/_next/image?url=%2Fmodels%2Flogo%2Fmind-lab-logo.png&w=48&q=75)Mind Lab

--

#### [Macaron V1 Venti](/models/model-detail/mindai-macaron-v1-venti?from=pricing)

[More](/models/model-detail/mindai-macaron-v1-venti?from=pricing)

Context1M

Input$1.5 /Mt· Cache Read $0.3 /Mt

Output$4.5 /Mt

#### [Macaron V1 Tall](/models/model-detail/mindai-macaron-v1-tall?from=pricing)

[More](/models/model-detail/mindai-macaron-v1-tall?from=pricing)

Context256K

Input$0.45 /Mt· Cache Read $0.08 /Mt

Output$2.6 /Mt

![mind lab](/_next/image?url=%2Fmodels%2Flogo%2Fmind-lab-logo.png&w=48&q=75)Mind Lab

--

Model NameContextInputOutputActions[Macaron V1 Venti](/models/model-detail/mindai-macaron-v1-venti?from=pricing)1M$1.5 /Mt· Cache Read $0.3 /Mt$4.5 /Mt[More](/models/model-detail/mindai-macaron-v1-venti?from=pricing)[Macaron V1 Tall](/models/model-detail/mindai-macaron-v1-tall?from=pricing)256K$0.45 /Mt· Cache Read $0.08 /Mt$2.6 /Mt[More](/models/model-detail/mindai-macaron-v1-tall?from=pricing)

![minimax](/models/logo/svg/minimax-logo.svg)MiniMax

Minimax AI's advanced language models delivering robust conversational AI capabilities with optimized performance for customer service, content generation, and creative applications, featuring strong multilingual support and enterprise-ready scalability.

#### [MiniMax M3](/models/model-detail/minimax-minimax-m3?from=pricing)

[More](/models/model-detail/minimax-minimax-m3?from=pricing)

Context977K

Input-

OutputTiered pricing

#### [MiniMax M2.7](/models/model-detail/minimax-minimax-m2.7?from=pricing)

[More](/models/model-detail/minimax-minimax-m2.7?from=pricing)

Context200K

Input$0.3 /Mt· Cache Read $0.06 /Mt

Output$1.2 /Mt

#### [MiniMax M2.5-highspeed](/models/model-detail/minimax-minimax-m2.5-highspeed?from=pricing)

[More](/models/model-detail/minimax-minimax-m2.5-highspeed?from=pricing)

Context200K

Input$0.6 /Mt· Cache Read $0.03 /Mt

Output$2.4 /Mt

#### [MiniMax M2.5](/models/model-detail/minimax-minimax-m2.5?from=pricing)

[More](/models/model-detail/minimax-minimax-m2.5?from=pricing)

Context200K

Input$0.3 /Mt· Cache Read $0.03 /Mt

Output$1.2 /Mt

#### [Minimax M2.1](/models/model-detail/minimax-minimax-m2.1?from=pricing)

[More](/models/model-detail/minimax-minimax-m2.1?from=pricing)

Context200K

Input$0.3 /Mt· Cache Read $0.03 /Mt

Output$1.2 /Mt

#### [MiniMax-M2](/models/model-detail/minimax-minimax-m2?from=pricing)

[More](/models/model-detail/minimax-minimax-m2?from=pricing)

Context200K

Input$0.3 /Mt· Cache Read $0.03 /Mt

Output$1.2 /Mt

#### [MiniMax M1](/models/model-detail/minimaxai-minimax-m1-80k?from=pricing)

[More](/models/model-detail/minimaxai-minimax-m1-80k?from=pricing)

Context977K

Input$0.55 /Mt

Output$2.2 /Mt

![minimax](/models/logo/svg/minimax-logo.svg)MiniMax

Minimax AI's advanced language models delivering robust conversational AI capabilities with optimized performance for customer service, content generation, and creative applications, featuring strong multilingual support and enterprise-ready scalability.

Model NameContextInputOutputActions[MiniMax M3](/models/model-detail/minimax-minimax-m3?from=pricing)977K-Tiered pricing[More](/models/model-detail/minimax-minimax-m3?from=pricing)[MiniMax M2.7](/models/model-detail/minimax-minimax-m2.7?from=pricing)200K$0.3 /Mt· Cache Read $0.06 /Mt$1.2 /Mt[More](/models/model-detail/minimax-minimax-m2.7?from=pricing)[MiniMax M2.5-highspeed](/models/model-detail/minimax-minimax-m2.5-highspeed?from=pricing)200K$0.6 /Mt· Cache Read $0.03 /Mt$2.4 /Mt[More](/models/model-detail/minimax-minimax-m2.5-highspeed?from=pricing)[MiniMax M2.5](/models/model-detail/minimax-minimax-m2.5?from=pricing)200K$0.3 /Mt· Cache Read $0.03 /Mt$1.2 /Mt[More](/models/model-detail/minimax-minimax-m2.5?from=pricing)[Minimax M2.1](/models/model-detail/minimax-minimax-m2.1?from=pricing)200K$0.3 /Mt· Cache Read $0.03 /Mt$1.2 /Mt[More](/models/model-detail/minimax-minimax-m2.1?from=pricing)[MiniMax-M2](/models/model-detail/minimax-minimax-m2?from=pricing)200K$0.3 /Mt· Cache Read $0.03 /Mt$1.2 /Mt[More](/models/model-detail/minimax-minimax-m2?from=pricing)[MiniMax M1](/models/model-detail/minimaxai-minimax-m1-80k?from=pricing)977K$0.55 /Mt$2.2 /Mt[More](/models/model-detail/minimaxai-minimax-m1-80k?from=pricing)

![inclusionai](/_next/image?url=%2Fmodels%2Flogo%2Finclusionai-logo.png&w=48&q=75)inclusionai

--

#### [Ling 3.0 Flash Fast](/models/model-detail/inclusionai-ling-3.0-flash-fast?from=pricing)

[More](/models/model-detail/inclusionai-ling-3.0-flash-fast?from=pricing)

Context256K

Input$0.06 /Mt· Cache Read $0.012 /Mt

Output$0.18 /Mt

#### [Ling 3.0 Flash](/models/model-detail/inclusionai-ling-3.0-flash?from=pricing)

[More](/models/model-detail/inclusionai-ling-3.0-flash?from=pricing)

Context256K

Input$0.06 /Mt· Cache Read $0.012 /Mt

Output$0.18 /Mt

#### [Ling-2.6-flash](/models/model-detail/inclusionai-ling-2.6-flash?from=pricing)

[More](/models/model-detail/inclusionai-ling-2.6-flash?from=pricing)

Context256K

Input$0.1 /Mt· Cache Read $0.02 /Mt

Output$0.3 /Mt

#### [Ling-2.6-1T](/models/model-detail/inclusionai-ling-2.6-1t?from=pricing)

[More](/models/model-detail/inclusionai-ling-2.6-1t?from=pricing)

Context256K

Input$0.3 /Mt· Cache Read $0.06 /Mt

Output$2.5 /Mt

![inclusionai](/_next/image?url=%2Fmodels%2Flogo%2Finclusionai-logo.png&w=48&q=75)inclusionai

--

Model NameContextInputOutputActions[Ling 3.0 Flash Fast](/models/model-detail/inclusionai-ling-3.0-flash-fast?from=pricing)256K$0.06 /Mt· Cache Read $0.012 /Mt$0.18 /Mt[More](/models/model-detail/inclusionai-ling-3.0-flash-fast?from=pricing)[Ling 3.0 Flash](/models/model-detail/inclusionai-ling-3.0-flash?from=pricing)256K$0.06 /Mt· Cache Read $0.012 /Mt$0.18 /Mt[More](/models/model-detail/inclusionai-ling-3.0-flash?from=pricing)[Ling-2.6-flash](/models/model-detail/inclusionai-ling-2.6-flash?from=pricing)256K$0.1 /Mt· Cache Read $0.02 /Mt$0.3 /Mt[More](/models/model-detail/inclusionai-ling-2.6-flash?from=pricing)[Ling-2.6-1T](/models/model-detail/inclusionai-ling-2.6-1t?from=pricing)256K$0.3 /Mt· Cache Read $0.06 /Mt$2.5 /Mt[More](/models/model-detail/inclusionai-ling-2.6-1t?from=pricing)

![stepfun](/models/logo/svg/stepfun-logo.svg)StepFun

--

#### [Step 3.7 Flash](/models/model-detail/stepfun-step-3.7-flash?from=pricing)

[More](/models/model-detail/stepfun-step-3.7-flash?from=pricing)

Context256K

Input$0.2 /Mt· Cache Read $0.04 /Mt

Output$1.15 /Mt

![stepfun](/models/logo/svg/stepfun-logo.svg)StepFun

--

Model NameContextInputOutputActions[Step 3.7 Flash](/models/model-detail/stepfun-step-3.7-flash?from=pricing)256K$0.2 /Mt· Cache Read $0.04 /Mt$1.15 /Mt[More](/models/model-detail/stepfun-step-3.7-flash?from=pricing)

![nvidia](/models/logo/svg/nvidia-logo.svg)Nvidia

--

#### [Nemotron 3 Nano 30B A3B](/models/model-detail/nvidia-nemotron-3-nano-30b-a3b?from=pricing)

[More](/models/model-detail/nvidia-nemotron-3-nano-30b-a3b?from=pricing)

Context256K

Input$0.05 /Mt

Output$0.2 /Mt

![nvidia](/models/logo/svg/nvidia-logo.svg)Nvidia

--

Model NameContextInputOutputActions[Nemotron 3 Nano 30B A3B](/models/model-detail/nvidia-nemotron-3-nano-30b-a3b?from=pricing)256K$0.05 /Mt$0.2 /Mt[More](/models/model-detail/nvidia-nemotron-3-nano-30b-a3b?from=pricing)

![gemma](/models/logo/svg/google-logo.svg)Gemma

Google's Gemma models offering high-quality language processing with excellent performance for various NLP tasks.

#### [Gemma 4 26B A4B](/models/model-detail/google-gemma-4-26b-a4b-it?from=pricing)

[More](/models/model-detail/google-gemma-4-26b-a4b-it?from=pricing)

Context256K

Input$0.13 /Mt

Output$0.4 /Mt

#### [Gemma 4 31B](/models/model-detail/google-gemma-4-31b-it?from=pricing)

[More](/models/model-detail/google-gemma-4-31b-it?from=pricing)

Context256K

Input$0.14 /Mt

Output$0.4 /Mt

#### [Gemma 3 27B](/models/model-detail/google-gemma-3-27b-it?from=pricing)

[More](/models/model-detail/google-gemma-3-27b-it?from=pricing)

Context96K

Input$0.119 /Mt

Output$0.2 /Mt

![gemma](/models/logo/svg/google-logo.svg)Gemma

Google's Gemma models offering high-quality language processing with excellent performance for various NLP tasks.

Model NameContextInputOutputActions[Gemma 4 26B A4B](/models/model-detail/google-gemma-4-26b-a4b-it?from=pricing)256K$0.13 /Mt$0.4 /Mt[More](/models/model-detail/google-gemma-4-26b-a4b-it?from=pricing)[Gemma 4 31B](/models/model-detail/google-gemma-4-31b-it?from=pricing)256K$0.14 /Mt$0.4 /Mt[More](/models/model-detail/google-gemma-4-31b-it?from=pricing)[Gemma 3 27B](/models/model-detail/google-gemma-3-27b-it?from=pricing)96K$0.119 /Mt$0.2 /Mt[More](/models/model-detail/google-gemma-3-27b-it?from=pricing)

![kwaikat](/models/logo/svg/StreamLake.svg)KwaiKAT

--

#### [Kat Coder Pro](/models/model-detail/kwaipilot-kat-coder-pro?from=pricing)

[More](/models/model-detail/kwaipilot-kat-coder-pro?from=pricing)

Context250K

Input$0.3 /Mt· Cache Read $0.06 /Mt

Output$1.2 /Mt

![kwaikat](/models/logo/svg/StreamLake.svg)KwaiKAT

--

Model NameContextInputOutputActions[Kat Coder Pro](/models/model-detail/kwaipilot-kat-coder-pro?from=pricing)250K$0.3 /Mt· Cache Read $0.06 /Mt$1.2 /Mt[More](/models/model-detail/kwaipilot-kat-coder-pro?from=pricing)

OpenAI

--

#### [OpenAI GPT OSS 120B](/models/model-detail/openai-gpt-oss-120b?from=pricing)

[More](/models/model-detail/openai-gpt-oss-120b?from=pricing)

Context128K

Input$0.05 /Mt

Output$0.25 /Mt

#### [OpenAI: GPT OSS 20B](/models/model-detail/openai-gpt-oss-20b?from=pricing)

[More](/models/model-detail/openai-gpt-oss-20b?from=pricing)

Context128K

Input$0.04 /Mt

Output$0.15 /Mt

OpenAI

--

Model NameContextInputOutputActions[OpenAI GPT OSS 120B](/models/model-detail/openai-gpt-oss-120b?from=pricing)128K$0.05 /Mt$0.25 /Mt[More](/models/model-detail/openai-gpt-oss-120b?from=pricing)[OpenAI: GPT OSS 20B](/models/model-detail/openai-gpt-oss-20b?from=pricing)128K$0.04 /Mt$0.15 /Mt[More](/models/model-detail/openai-gpt-oss-20b?from=pricing)

![llama](/models/logo/svg/meta-logo.svg)Llama

Meta's Llama models providing state-of-the-art language understanding with open architecture designed for diverse applications.

#### [Llama 3.1 8B Instruct](/models/model-detail/meta-llama-llama-3.1-8b-instruct?from=pricing)

[More](/models/model-detail/meta-llama-llama-3.1-8b-instruct?from=pricing)

Context16K

Input$0.02 /Mt

Output$0.05 /Mt

#### [Llama 3.3 70B Instruct](/models/model-detail/meta-llama-llama-3.3-70b-instruct?from=pricing)

[More](/models/model-detail/meta-llama-llama-3.3-70b-instruct?from=pricing)

Context12K

Input$0.135 /Mt

Output$0.4 /Mt

#### [Llama 4 Maverick Instruct](/models/model-detail/meta-llama-llama-4-maverick-17b-128e-instruct-fp8?from=pricing)

[More](/models/model-detail/meta-llama-llama-4-maverick-17b-128e-instruct-fp8?from=pricing)

Context1M

Input$0.27 /Mt

Output$0.85 /Mt

#### [Llama 4 Scout Instruct](/models/model-detail/meta-llama-llama-4-scout-17b-16e-instruct?from=pricing)

[More](/models/model-detail/meta-llama-llama-4-scout-17b-16e-instruct?from=pricing)

Context128K

Input$0.18 /Mt

Output$0.59 /Mt

![llama](/models/logo/svg/meta-logo.svg)Llama

Meta's Llama models providing state-of-the-art language understanding with open architecture designed for diverse applications.

Model NameContextInputOutputActions[Llama 3.1 8B Instruct](/models/model-detail/meta-llama-llama-3.1-8b-instruct?from=pricing)16K$0.02 /Mt$0.05 /Mt[More](/models/model-detail/meta-llama-llama-3.1-8b-instruct?from=pricing)[Llama 3.3 70B Instruct](/models/model-detail/meta-llama-llama-3.3-70b-instruct?from=pricing)12K$0.135 /Mt$0.4 /Mt[More](/models/model-detail/meta-llama-llama-3.3-70b-instruct?from=pricing)[Llama 4 Maverick Instruct](/models/model-detail/meta-llama-llama-4-maverick-17b-128e-instruct-fp8?from=pricing)1M$0.27 /Mt$0.85 /Mt[More](/models/model-detail/meta-llama-llama-4-maverick-17b-128e-instruct-fp8?from=pricing)[Llama 4 Scout Instruct](/models/model-detail/meta-llama-llama-4-scout-17b-16e-instruct?from=pricing)128K$0.18 /Mt$0.59 /Mt[More](/models/model-detail/meta-llama-llama-4-scout-17b-16e-instruct?from=pricing)

Mistral

--

#### [Mistral Nemo](/models/model-detail/mistralai-mistral-nemo?from=pricing)

[More](/models/model-detail/mistralai-mistral-nemo?from=pricing)

Context59K

Input$0.04 /Mt

Output$0.17 /Mt

Mistral

--

Model NameContextInputOutputActions[Mistral Nemo](/models/model-detail/mistralai-mistral-nemo?from=pricing)59K$0.04 /Mt$0.17 /Mt[More](/models/model-detail/mistralai-mistral-nemo?from=pricing)

OOthers

--

#### [XiaomiMiMo/MiMo-V2.5](/models/model-detail/xiaomimimo-mimo-v2.5?from=pricing)

[More](/models/model-detail/xiaomimimo-mimo-v2.5?from=pricing)

Context1M

Input$0.168 /Mt· Cache Read $0.0034 /Mt

Output$0.336 /Mt

#### [XiaomiMiMo/MiMo-V2.5-Pro](/models/model-detail/xiaomimimo-mimo-v2.5-pro?from=pricing)

[More](/models/model-detail/xiaomimimo-mimo-v2.5-pro?from=pricing)

Context1M

Input$0.522 /Mt· Cache Read $0.0043 /Mt

Output$1.044 /Mt

#### [Wizardlm 2 8x22B](/models/model-detail/microsoft-wizardlm-2-8x22b?from=pricing)

[More](/models/model-detail/microsoft-wizardlm-2-8x22b?from=pricing)

Context64K

Input$0.62 /Mt

Output$0.62 /Mt

#### [Ring-2.6-1T](/models/model-detail/inclusionai-ring-2.6-1t?from=pricing)

[More](/models/model-detail/inclusionai-ring-2.6-1t?from=pricing)

Context256K

Input$0.3 /Mt· Cache Read $0.06 /Mt

Output$2.5 /Mt

OOthers

--

Model NameContextInputOutputActions[XiaomiMiMo/MiMo-V2.5](/models/model-detail/xiaomimimo-mimo-v2.5?from=pricing)1M$0.168 /Mt· Cache Read $0.0034 /Mt$0.336 /Mt[More](/models/model-detail/xiaomimimo-mimo-v2.5?from=pricing)[XiaomiMiMo/MiMo-V2.5-Pro](/models/model-detail/xiaomimimo-mimo-v2.5-pro?from=pricing)1M$0.522 /Mt· Cache Read $0.0043 /Mt$1.044 /Mt[More](/models/model-detail/xiaomimimo-mimo-v2.5-pro?from=pricing)[Wizardlm 2 8x22B](/models/model-detail/microsoft-wizardlm-2-8x22b?from=pricing)64K$0.62 /Mt$0.62 /Mt[More](/models/model-detail/microsoft-wizardlm-2-8x22b?from=pricing)[Ring-2.6-1T](/models/model-detail/inclusionai-ring-2.6-1t?from=pricing)256K$0.3 /Mt· Cache Read $0.06 /Mt$2.5 /Mt[More](/models/model-detail/inclusionai-ring-2.6-1t?from=pricing)

### Embeddings

#### qwen/qwen3-embedding-0.6b

Context32K

Input$0.07 /Mt

#### Qwen3 Embedding 8B

Context32K

Input$0.07 /Mt

#### BAAI:BGE-M3

Context8K

Input$0.01 /Mt

Model NameContextInput

qwen/qwen3-embedding-0.6b

32K$0.07 /Mt

Qwen3 Embedding 8B

32K$0.07 /Mt

BAAI:BGE-M3

8K$0.01 /Mt

### Image

Pricing may vary based on image dimensions, inference steps, and upscaling factors. Use thePricing Calculatorfor an estimate.

#### Flux.1 Kontext Dev

Mode-

Width&Height-

Pricing$0.0225 /image

#### flux1-kontext-dev

Modefast_mode

Width&Height-

Pricing$0.018 /image

#### Flux.1 Kontext Max

Mode-

Width&Height-

Pricing$0.072 /image

#### Flux.1 Kontext Pro

Mode-

Width&Height-

Pricing$0.36 /image

#### Qwen-Image Edit

Mode-

Width&Height-

Pricing$0.02 /image

#### Qwen-Image Text to Image

Mode-

Width&Height-

Pricing$0.02 /image

API NameModeWidth&HeightPricing

Flux.1 Kontext Dev

--$0.0225 /imagefast_mode-$0.018 /image

Flux.1 Kontext Max

--$0.072 /image

Flux.1 Kontext Pro

--$0.36 /image

Qwen-Image Edit

--$0.02 /image

Qwen-Image Text to Image

--$0.02 /image

### Video

Pricing may vary based on the number of frames, chosen model, and inference steps. Use thePricing Calculatorfor an estimate.

#### Kling v3.0 Pro Image-to-Video

ModeNo Audio

Duration-

Resolution

Pricing$0.112 /s

#### kling-v3.0-pro-i2v

ModeAudio

Duration-

Resolution

Pricing$0.168 /s

#### Kling v3.0 Pro Text-to-Video

ModeNo Audio

Duration-

Resolution

Pricing$0.112 /s

#### kling-v3.0-pro-t2v

ModeAudio

Duration-

Resolution

Pricing$0.168 /s

#### Kling v3.0 Standard Image-to-Video

ModeNo Audio

Duration-

Resolution

Pricing$0.084 /s

#### kling-v3.0-std-i2v

ModeAudio

Duration-

Resolution

Pricing$0.126 /s

#### Kling v3.0 Standard Text-to-Video

ModeNo Audio

Duration-

Resolution

Pricing$0.084 /s

#### kling-v3.0-std-t2v

ModeAudio

Duration-

Resolution

Pricing$0.126 /s

#### Minimax Hailuo 2.3 Fast Image to Video

Mode-

Duration6s

Resolution768P

Pricing$0.19 /video

#### minimax-hailuo-2.3-fast-i2v

Mode-

Duration10s

Resolution768P

Pricing$0.32 /video

#### minimax-hailuo-2.3-fast-i2v

Mode-

Duration6s

Resolution1080P

Pricing$0.33 /video

#### Minimax Hailuo 2.3 Image to Video

Mode-

Duration6s

Resolution768P

Pricing$0.28 /video

#### minimax-hailuo-2.3-i2v

Mode-

Duration10s

Resolution768P

Pricing$0.56 /video

#### minimax-hailuo-2.3-i2v

Mode-

Duration6s

Resolution1080P

Pricing$0.49 /video

#### Minimax Hailuo 2.3 Text to Video

Mode-

Duration6s

Resolution768P

Pricing$0.28 /video

#### minimax-hailuo-2.3-t2v

Mode-

Duration10s

Resolution768P

Pricing$0.56 /video

#### minimax-hailuo-2.3-t2v

Mode-

Duration6s

Resolution1080P

Pricing$0.49 /video

#### Wan 2.5 Image to Video

Mode-

Duration5s

Resolution480P

Pricing$0.25 /video

#### wan-2.5-i2v

Mode-

Duration10s

Resolution480P

Pricing$0.50 /video

#### wan-2.5-i2v

Mode-

Duration5s

Resolution720P

Pricing$0.50 /video

#### wan-2.5-i2v

Mode-

Duration10s

Resolution720P

Pricing$1.00 /video

#### wan-2.5-i2v

Mode-

Duration5s

Resolution1080P

Pricing$0.75 /video

#### wan-2.5-i2v

Mode-

Duration10s

Resolution1080P

Pricing$1.50 /video

#### Wan 2.5 Text to Video

Mode-

Duration5s

Resolution480P

Pricing$0.25 /video

#### wan-2.5-t2v

Mode-

Duration10s

Resolution480P

Pricing$0.50 /video

#### wan-2.5-t2v

Mode-

Duration5s

Resolution720P

Pricing$0.50 /video

#### wan-2.5-t2v

Mode-

Duration10s

Resolution720P

Pricing$1.00 /video

#### wan-2.5-t2v

Mode-

Duration5s

Resolution1080P

Pricing$0.75 /video

#### wan-2.5-t2v

Mode-

Duration10s

Resolution1080P

Pricing$1.50 /video

#### Wan 2.6 Image to Video

Mode-

Duration5s

Resolution720P

Pricing$0.50 /video

#### wan-2.6-i2v

Mode-

Duration10s

Resolution720P

Pricing$1.00 /video

#### wan-2.6-i2v

Mode-

Duration15s

Resolution720P

Pricing$1.50 /video

#### wan-2.6-i2v

Mode-

Duration5s

Resolution1080P

Pricing$0.75 /video

#### wan-2.6-i2v

Mode-

Duration10s

Resolution1080P

Pricing$1.50 /video

#### wan-2.6-i2v

Mode-

Duration15s

Resolution1080P

Pricing$2.25 /video

#### Wan 2.6 Reference to Video

Mode-

Duration5s

Resolution720P

Pricing$0.50 /video

#### wan-2.6-v2v

Mode-

Duration10s

Resolution720P

Pricing$1.00 /video

#### wan-2.6-v2v

Mode-

Duration5s

Resolution1080P

Pricing$0.75 /video

#### wan-2.6-v2v

Mode-

Duration10s

Resolution1080P

Pricing$1.50 /video

#### Wan 2.6 Text to Video

Mode-

Duration5s

Resolution720P

Pricing$0.50 /video

#### wan-2.6-t2v

Mode-

Duration10s

Resolution720P

Pricing$1.00 /video

#### wan-2.6-t2v

Mode-

Duration15s

Resolution720P

Pricing$1.50 /video

#### wan-2.6-t2v

Mode-

Duration5s

Resolution1080P

Pricing$0.75 /video

#### wan-2.6-t2v

Mode-

Duration10s

Resolution1080P

Pricing$1.50 /video

#### wan-2.6-t2v

Mode-

Duration15s

Resolution1080P

Pricing$2.25 /video

#### Image to Video

ModelSVD-XT

Steps20

Pricing$0.024 /video

#### img2video

ModelSVD

Steps20

Pricing$0.0134 /video

API NameModeDurationResolutionPricing

Kling v3.0 Pro Image-to-Video

No Audio-$0.112 /sAudio-$0.168 /s

Kling v3.0 Pro Text-to-Video

No Audio-$0.112 /sAudio-$0.168 /s

Kling v3.0 Standard Image-to-Video

No Audio-$0.084 /sAudio-$0.126 /s

Kling v3.0 Standard Text-to-Video

No Audio-$0.084 /sAudio-$0.126 /s

Minimax Hailuo 2.3 Fast Image to Video

-6s768P$0.19 /video-10s768P$0.32 /video-6s1080P$0.33 /video

Minimax Hailuo 2.3 Image to Video

-6s768P$0.28 /video-10s768P$0.56 /video-6s1080P$0.49 /video

Minimax Hailuo 2.3 Text to Video

-6s768P$0.28 /video-10s768P$0.56 /video-6s1080P$0.49 /video

Wan 2.5 Image to Video

-5s480P$0.25 /video-10s480P$0.50 /video-5s720P$0.50 /video-10s720P$1.00 /video-5s1080P$0.75 /video-10s1080P$1.50 /video

Wan 2.5 Text to Video

-5s480P$0.25 /video-10s480P$0.50 /video-5s720P$0.50 /video-10s720P$1.00 /video-5s1080P$0.75 /video-10s1080P$1.50 /video

Wan 2.6 Image to Video

-5s720P$0.50 /video-10s720P$1.00 /video-15s720P$1.50 /video-5s1080P$0.75 /video-10s1080P$1.50 /video-15s1080P$2.25 /video

Wan 2.6 Reference to Video

-5s720P$0.50 /video-10s720P$1.00 /video-5s1080P$0.75 /video-10s1080P$1.50 /video

Wan 2.6 Text to Video

-5s720P$0.50 /video-10s720P$1.00 /video-15s720P$1.50 /video-5s1080P$0.75 /video-10s1080P$1.50 /video-15s1080P$2.25 /video

API NameModelStepsPricing

Image to Video

SVD-XT20$0.024 /videoSVD20$0.0134 /video

### Audio

#### Fish Audio Text to Speech

Mode-

Pricing$15 /1M characters

#### Fish Audio Voice Cloning

Mode-

Pricing$0.1 /voice

#### MiniMax speech-2.6-hd

ModeT2A / T2A Async

Pricing$100 /1M characters

#### MiniMax speech-2.6-turbo

ModeT2A / T2A Async

Pricing$60 /1M characters

#### MiniMax Voice-Cloning

Mode-

Pricing$1.5 /voice

#### Text to Speech

Mode-

Pricing$15 /1M characters

API NameModePricing

Fish Audio Text to Speech

-$15 /1M characters

Fish Audio Voice Cloning

-$0.1 /voice

MiniMax speech-2.6-hd

T2A / T2A Async$100 /1M characters

MiniMax speech-2.6-turbo

T2A / T2A Async$60 /1M characters

MiniMax Voice-Cloning

-$1.5 /voice

Text to Speech

-$15 /1M characters

### AI Search

#### neuralSearch

API NameEXA

Pricing$0.007/request

#### deepSearch

API NameEXA

Pricing$0.012/request

#### deepReasoningSearch

API NameEXA

Pricing$0.015/request

#### additional_result

API NameEXA

Pricing$0.001/result (beyond 10 results)

#### answer

API NameEXA

Pricing$0.005/request

#### contentText

API NameEXA

Pricing$0.001/item

#### contentHighlight

API NameEXA

Pricing$0.001/item

#### contentSummary

API NameEXA

Pricing$0.001/item

#### basicSearch

API NameTavily

Pricing$0.008/request

#### advancedSearch

API NameTavily

Pricing$0.016/request

#### basicExtract

API NameTavily

Pricing$0.0016/url

#### advancedExtract

API NameTavily

Pricing$0.0032/url

#### regularMapping

API NameTavily

Pricing$0.0008/page

#### instructedMapping

API NameTavily

Pricing$0.0016/page

#### Crawl

API NameTavily

PricingExtract + Mapping Cost

API NameModePricingEXAneuralSearch$0.007/requestdeepSearch$0.012/requestdeepReasoningSearch$0.015/requestadditional_result$0.001/result (beyond 10 results)answer$0.005/requestcontentText$0.001/itemcontentHighlight$0.001/itemcontentSummary$0.001/itemTavilybasicSearch$0.008/requestadvancedSearch$0.016/requestbasicExtract$0.0016/urladvancedExtract$0.0032/urlregularMapping$0.0008/pageinstructedMapping$0.0016/pageCrawlExtract + Mapping Cost

## Everything you need to build production AI.

200+ models, on-demand GPUs, and secure agent runtimes — unified under one API. Free to start, scales as you grow.

Get Started

Power your AI applications with Novita AI's model APIs, GPU instances, and agent sandbox.

![AICPA SOC 2 Certified](/_next/image?url=%2Ffooter%2Fv5%2Fgtrp.png&w=128&q=75)

Product

[Model APIs](/models)

[Agent Sandbox](/sandbox)

[GPU Instance](/gpus)

[GPU Bare Metal](/gpu-baremetal)

Resource

[Talk to Sales](https://meetings-na2.hubspot.com/junyu)

[Contact Support](mailto:support@novita.ai)

[Pricing](/pricing)

[Docs](/docs/guides/introduction)

Partners

[Refer to Earn](/affiliate-new)

[Supply GPUs](mailto:gpu@novita.ai)

Company

[Careers](https://jobs.ashbyhq.com/novita-ai)

[Blog](https://blogs.novita.ai)

[Trust Center](https://trust.novita.ai)

© 2026 Novita AI. All rights reserved

[Terms of Service](/legal/terms-of-service)[Privacy Policy](/legal/privacy-policy)[Cookie Policy](/legal/cookie-policy)Cookie SettingsDo Not Sell or Share My Personal Information
