OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Llama 3.1 (405B) Instruct Turbo This page | ||||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
Yes, Llama 3.1 (405B) Instruct Turbo can stream responses as they are generated.
Yes, Llama 3.1 (405B) Instruct Turbo supports both function calling and structured outputs.
Llama 3.1 (405B) Instruct Turbo was built by Meta.
The model is available through the v1/chat/completions endpoint.