Llama 3.1 (405B) Instruct Turbo API

Llama 3.1 405B: Advanced text generation model by Meta AI.
Output

How to use Llama 3.1 (405B) Instruct Turbo API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to .

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Llama 3.1 (405B) Instruct Turbo API Pricing

TypePrice
Input
Output

Llama 3.1 (405B) Instruct Turbo vs other models

ModelInputOutputContextBest for
$6.5 / 1M tokens
$39 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1.05M tokens
Reasoning + agents

Frequently asked questions

Yes, Llama 3.1 (405B) Instruct Turbo can stream responses as they are generated.

Yes, Llama 3.1 (405B) Instruct Turbo supports both function calling and structured outputs.

Llama 3.1 (405B) Instruct Turbo was built by Meta.

The model is available through the v1/chat/completions endpoint.

Start building with Llama 3.1 (405B) Instruct Turbo

Get API Key
1000+ models, one API.