OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Llama 3.2 3B Instruct Turbo This page | ||||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
Yes, Llama 3.2 3B Instruct Turbo can stream responses as they are generated.
Yes, Llama 3.2 3B Instruct Turbo supports both function calling and structured outputs.
Llama 3.2 3B Instruct Turbo was built by Meta.
It is available at the v1/chat/completions endpoint.