import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "meta-llama/Llama-3.3-70B-Instruct-Turbo", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "meta-llama/Llama-3.3-70B-Instruct-Turbo", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"meta-llama/Llama-3.3-70B-Instruct-Turbo","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Llama 3.3 70B Instruct Turbo This page | Coding + agents | |||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
Llama 3.3 70B Instruct Turbo has a 128,000 tokens context window and can return up to 127,000 tokens.
Llama 3.3 70B Instruct Turbo takes text as input and returns text.
Use meta-llama/Llama-3.3-70B-Instruct-Turbo as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.
Llama 3.3 70B Instruct Turbo became available on September 30, 2025.
Llama 3.3 70B Instruct Turbo is priced at input $1.144 / 1M tokens, output $1.144 / 1M tokens.
Yes, Llama 3.3 70B Instruct Turbo can stream responses as they are generated.
Llama 3.3 70B Instruct Turbo was built by Meta.
Yes, it supports function calling, parallel tool calls, and tool integration.