import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "alibaba/qwen3.8-max-prime", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "alibaba/qwen3.8-max-prime", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"alibaba/qwen3.8-max-prime","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
Prime is the higher-throughput SKU and is priced above the standard Qwen3.8 Max ($2.6 / $7.8 per 1M tokens).
Throughput-bound agent loops. Pipelines where wall-clock time per step decides whether the workflow is usable. Prime is the same model family served for higher throughput, which is the only reason to pay its premium over the standard SKU.
Video and image analysis at volume. Video joins text and images as an input modality, so clips are analysed in the same request as the prompt rather than through a separate pipeline.
Repository-scale code work. A 1,000,000-token window holds a large codebase with its tests, and 131,072 tokens of output is room for a substantial refactor in one reply.
Long-document research. Contracts, filings and archives handed over whole, with always-on reasoning applied to every request.
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Qwen3.8 Max Prime This page | Throughput-bound multimodal reasoning | |||
| Complex reasoning, coding and agentic workflows | ||||
| Frontier reasoning + agents | ||||
| Reasoning + agents | ||||
| Long-context, multimodal & agentic workflows |
Qwen3.8 Max Prime is a higher-throughput variant of Alibaba's flagship Qwen3.8 Max, served as its own SKU. It keeps the 1,000,000-token context window, accepts text, image and video input, and always reasons.
Prime is served for higher throughput and adds video as an input modality, and its reasoning is always on rather than switchable. Both share the same 1M-token context and 131,072-token output limit. Prime costs more: $5.5016 / $16.5048 per 1M tokens against $2.6 / $7.8 for the standard SKU.
On AI/ML API it costs $5.5016 per 1M input tokens and $16.5048 per 1M output tokens, with cached input at $0.6877 per 1M. Reasoning tokens are billed at the output rate.
Text, images and video in, text out, with up to 1,000,000 tokens in a single request and up to 131,072 tokens of output.
Send a POST request to https://api.aimlapi.com/v1/chat/completions with alibaba/qwen3.8-max-prime as the model id and your AI/ML API key in the Authorization header. The endpoint is OpenAI-compatible, so an existing chat client only needs the base URL swapped.