MODEL APIS

Browse our supported open source models

Developer-first infrastructure that scales from zero to production.

Models (15)
DeepSeekDeepSeek
New
DeepSeek V4.1 FlashDeepSeek V4.1 Flash
$0.3/MtInput
$0.006/MtCache Read
$1.2/MtOutput
1MContext
384KMax Output
LLMServerless
New
DeepSeek V4 Pro 0813DeepSeek V4 Pro 0813
$1.32/MtInput
$0.044/MtCache Read
$3.96/MtOutput
1MContext
384KMax Output
LLMServerless
New
DeepSeek V4 Flash Vision ExpDeepSeek V4 Flash Vision Exp
$0.44/MtInput
$0.028/MtCache Read
$1.32/MtOutput
1MContext
384KMax Output
LLMServerless
New
DeepSeek V4 Flash 0731DeepSeek V4 Flash 0731
$0.44/MtInput
$0.028/MtCache Read
$1.32/MtOutput
1MContext
384KMax Output
LLMServerless
New
DeepSeek V4 FlashDeepSeek V4 Flash
$0.14/MtInput
$0.028/MtCache Read
$0.28/MtOutput
1MContext
384KMax Output
LLMServerless
New
DeepSeek V4 ProDeepSeek V4 Pro
$1.6/MtInput
$0.135/MtCache Read
$3.2/MtOutput
1MContext
384KMax Output
LLMServerless
Hot
DeepSeek V3.2DeepSeek V3.2
$0.269/MtInput
$0.1345/MtCache Read
$0.4/MtOutput
160KContext
64KMax Output
LLMServerless
DeepSeek OCR 2DeepSeek OCR 2
$0.03/MtInput
$0.03/MtOutput
8KContext
8KMax Output
LLMServerless
DeepSeek V3.2 ExpDeepSeek V3.2 Exp
$0.27/MtInput
$0.41/MtOutput
160KContext
64KMax Output
LLMServerless
DeepSeek V3.1 TerminusDeepSeek V3.1 Terminus
$0.27/MtInput
$0.135/MtCache Read
$1/MtOutput
128KContext
32KMax Output
LLMServerless
DeepSeek V3.1DeepSeek V3.1
$0.27/MtInput
$0.135/MtCache Read
$1/MtOutput
128KContext
32KMax Output
LLMServerless
DeepSeek R1 0528DeepSeek R1 0528
$0.7/MtInput
$0.35/MtCache Read
$2.5/MtOutput
160KContext
32KMax Output
LLMServerless
Dedicated
DeepSeek R1 Distill Qwen3 8B 0528DeepSeek R1 Distill Qwen3 8B 0528
$0.06/MtInput
$0.09/MtOutput
125KContext
31KMax Output
LLM
DeepSeek R1 Distill Llama 70BDeepSeek R1 Distill Llama 70B
$0.8/MtInput
$0.8/MtOutput
8KContext
8KMax Output
LLMServerless
DeepSeek R1 TurboDeepSeek R1 Turbo
$0.7/MtInput
$2.5/MtOutput
63KContext
16KMax Output
LLMServerless

Everything you need to build production AI.

200+ models, on-demand GPUs, and secure agent runtimes — unified under one API. Free to start, scales as you grow.