import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "inclusionai/ling-3.1-flash", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "inclusionai/ling-3.1-flash", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"inclusionai/ling-3.1-flash","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Output | |
No per-token charge is listed for this model in the AI/ML API model feed at the moment.
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| skillsBench | 68.7% | Breadth of discrete professional skills carried out end to end | Source | October 2, 2026 |
| AutomationBench | 52.5% | Completion of routine automation workflows | Source | October 2, 2026 |
| CyberGym | 87.9% | Security task completion in sandboxed environments | Source | October 2, 2026 |
| Finance Agent v2 | 57.9% | Financial research and analysis run as an agent | Source | October 2, 2026 |
| DRACO | 85.5% | Data research and analysis with complex operations | Source | October 2, 2026 |
| Terminal-Bench 4 | 40.4% | Autonomous shell and terminal task completion, fourth release | Source | October 2, 2026 |
| SWE-Atlas Codebase QnA | 55.9% | Answering questions about an unfamiliar codebase | Source | October 2, 2026 |
| HealthBench Professional | 65.3% | Clinical and professional health question answering | Source | October 2, 2026 |
Figures reported by inclusionAI at launch; not independently reproduced.
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Ling 3.1 Flash This page | Low-latency assistance, extraction and routine automation with tool calling | |||
| Fast, low-cost reasoning and tool use at scale | ||||
| Low-cost image and video understanding | ||||
| Reasoning + agents |
Yes. The model was announced by inclusionAI on 30 September 2026 and is available through AI/ML API's chat completions endpoint as inclusionai/ling-3.1-flash. The model feed currently lists no per-token charge for it.
It is a mixture-of-experts model with roughly 560 billion total parameters, of which about 25 billion are active for any given token. That ratio is the point of the design: the model has the capacity of a very large network while only paying to run a small slice of it per token, which is what keeps latency low.
256K tokens as served at launch, with a maximum output of 32,768 tokens. inclusionAI has said it targets a context window of up to one million tokens; until that is actually served, 256K is the figure to plan against.
inclusionAI positions it for low-latency assistance, extraction and routine automation. Its published figures are mostly agentic rather than conversational — skill execution, automation, terminal work and codebase question answering, plus domain suites in security, finance and health — so it is aimed at pipelines that hand the model a toolset and let it work through a task.
No. Every figure on this page was reported by inclusionAI at launch and none has been reproduced by a third party. The model's weights are not published either, so independent evaluation is not yet possible. Treat the numbers as the vendor's own claims.
Ling 3.0 Flash is a far smaller mixture-of-experts model — roughly 124 billion total parameters with about 5 billion active — and it is available today. Ling 3.1 Flash raises both figures by roughly four to five times, to about 560 billion total and 25 billion active, and adds hybrid reasoning. Both are served with a 256K context window.