Model Library/Llama 3.3 70B Instruct

Llama 3.3 70B Instruct

meta-llama/llama-3.3-70b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks. Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.

Características

API serverless

Documentación

meta-llama/llama-3.3-70b-instruct is available via Novita's serverless API, where you pay per token. There are several ways to call the API, including OpenAI-compatible endpoints with exceptional reasoning performance.

Implementaciones bajo demanda

Documentación

On-demand deployments allow you to use meta-llama/llama-3.3-70b-instruct on dedicated GPUs with high-performance serving stack with high reliability and no rate limits.

Serverless disponible

Ejecuta consultas de inmediato, paga solo por el uso

Entrada$0.135 / M Tokens

Salida$0.4 / M Tokens

Usa los siguientes ejemplos de código para integrarte con nuestra API:

1from openai import OpenAI
2
3client = OpenAI(
4    api_key="<Your API Key>",
5    base_url="https://api.novita.ai/openai"
6)
7
8response = client.chat.completions.create(
9    model="meta-llama/llama-3.3-70b-instruct",
10    messages=[
11        {"role": "system", "content": "You are a helpful assistant."},
12        {"role": "user", "content": "Hello, how are you?"}
13    ],
14    max_tokens=120000,
15    temperature=0.7
16)
17
18print(response.choices[0].message.content)

Información

Proveedor

Llama

Cuantización

bf16

Funcionalidad compatible

Longitud del contexto

6000

Salida máxima

120000

Serverless

Compatible

Function Calling

Compatible

Structured Output

Compatible

Capacidades de entrada

text

Capacidades de salida

text

Todo lo que necesitas para crear IA de producción.

Más de 200 modelos, GPUs bajo demanda y entornos de ejecución seguros para agentes, unificados bajo una API. Gratis para empezar, escala a medida que creces.