import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "inception/mercury-2.5", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "inception/mercury-2.5", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"inception/mercury-2.5","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Mercury 2.5 This page | Low-latency reasoning and agent loops | |||
| Reasoning + agents | ||||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Reasoning + agents |
Mercury 2.5 is a diffusion language model from Inception. Unlike models that emit one token after another, it generates and refines many tokens in parallel.
A standard LLM predicts tokens one at a time, each conditioned on the previous one. A diffusion language model starts from a rough draft of the whole output and refines it over several passes, so many tokens improve at once.
Inception reports throughput of 1,107 tokens per second on widely available NVIDIA GPUs, and describes it as the fastest reasoning model in production.
Yes. It supports function calling and can issue parallel tool calls within a single turn.
Yes. It supports schema-aligned JSON output, so responses can be constrained to the shape your code expects.
Yes. Reasoning depth is tunable, so you can trade latency against deliberation on a per-request basis.
Mercury 2.5 takes text and image input and returns text. It also supports file input and web search.
Send a request to the chat completions endpoint with the model id inception/mercury-2.5. The request shape is the same as for any other chat model on AI/ML API.
Inception reports roughly a 40% increase in intelligence over Mercury 2, while keeping the parallel decoding that makes the family fast.
Latency-sensitive work such as agent loops that make many sequential calls, and high-volume generation where throughput is the constraint.