Llama 3.3 70B Instruct Turbo API

meta-llama/Llama-3.3-70B-Instruct-Turbo
Meta Llama 3.3 70B Instruct Turbo is an advanced language model optimized for instruction-following tasks with high efficiency and performance.
Context
128K tokens
Input
$1.144 / 1M tokens
Output
$1.144 / 1M tokens
Released
Sep 30, 2025

How to use Llama 3.3 70B Instruct Turbo API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to meta-llama/Llama-3.3-70B-Instruct-Turbo.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "meta-llama/Llama-3.3-70B-Instruct-Turbo",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "meta-llama/Llama-3.3-70B-Instruct-Turbo",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"meta-llama/Llama-3.3-70B-Instruct-Turbo","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Llama 3.3 70B Instruct Turbo API Pricing

TypePrice
Input
$1.144 / 1M tokens
Output
$1.144 / 1M tokens

Llama 3.3 70B Instruct Turbo Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
GPQA Diamond
50.5%
Google-proof graduate science questions (hardest subset)SourceJuly 12, 2026
MMLU-Pro
68.9%
Multi-discipline knowledge + reasoning (harder MMLU)SourceJuly 12, 2026

Llama 3.3 70B Instruct Turbo vs other models

ModelInputOutputContextBest for
$1.144 / 1M tokens
$1.144 / 1M tokens
128K tokens
Coding + agents
$6.5 / 1M tokens
$39 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1.05M tokens
Reasoning + agents

Frequently asked questions

Llama 3.3 70B Instruct Turbo has a 128,000 tokens context window and can return up to 127,000 tokens.

Llama 3.3 70B Instruct Turbo takes text as input and returns text.

Use meta-llama/Llama-3.3-70B-Instruct-Turbo as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Llama 3.3 70B Instruct Turbo became available on September 30, 2025.

Llama 3.3 70B Instruct Turbo is priced at input $1.144 / 1M tokens, output $1.144 / 1M tokens.

Yes, Llama 3.3 70B Instruct Turbo can stream responses as they are generated.

Llama 3.3 70B Instruct Turbo was built by Meta.

Yes, it supports function calling, parallel tool calls, and tool integration.

Start building with Llama 3.3 70B Instruct Turbo

Get API Key
1000+ models, one API.