Qwen3.8 Max Prime API

alibaba/qwen3.8-max-prime
Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max, served as a separate SKU. It accepts text, image and video input over a 1M-token context, and reasoning is always on.
Context
1M tokens
Input
$5.5016 / 1M tokens
Output
$16.5048 / 1M tokens
Released
Sep 23, 2026

How to use Qwen3.8 Max Prime API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to alibaba/qwen3.8-max-prime.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "alibaba/qwen3.8-max-prime",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "alibaba/qwen3.8-max-prime",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"alibaba/qwen3.8-max-prime","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Qwen3.8 Max Prime API Pricing

TypePrice
Input
$5.5016 / 1M tokens
Output
$16.5048 / 1M tokens
Cached input
$0.6877 / 1M tokens

Prime is the higher-throughput SKU and is priced above the standard Qwen3.8 Max ($2.6 / $7.8 per 1M tokens).

Use cases

Throughput-bound agent loops. Pipelines where wall-clock time per step decides whether the workflow is usable. Prime is the same model family served for higher throughput, which is the only reason to pay its premium over the standard SKU.

Video and image analysis at volume. Video joins text and images as an input modality, so clips are analysed in the same request as the prompt rather than through a separate pipeline.

Repository-scale code work. A 1,000,000-token window holds a large codebase with its tests, and 131,072 tokens of output is room for a substantial refactor in one reply.

Long-document research. Contracts, filings and archives handed over whole, with always-on reasoning applied to every request.

Qwen3.8 Max Prime vs other models

ModelInputOutputContextBest for
$5.5016 / 1M tokens
$16.5048 / 1M tokens
1M tokens
Throughput-bound multimodal reasoning
$2.6 / 1M tokens
$7.8 / 1M tokens
1M tokens
Complex reasoning, coding and agentic workflows
$5.2 / 1M tokens
$26 / 1M tokens
1M tokens
Frontier reasoning + agents
$1.95 / 1M tokens
$9.75 / 1M tokens
1M tokens
Reasoning + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows

Frequently asked questions

Qwen3.8 Max Prime is a higher-throughput variant of Alibaba's flagship Qwen3.8 Max, served as its own SKU. It keeps the 1,000,000-token context window, accepts text, image and video input, and always reasons.

Prime is served for higher throughput and adds video as an input modality, and its reasoning is always on rather than switchable. Both share the same 1M-token context and 131,072-token output limit. Prime costs more: $5.5016 / $16.5048 per 1M tokens against $2.6 / $7.8 for the standard SKU.

On AI/ML API it costs $5.5016 per 1M input tokens and $16.5048 per 1M output tokens, with cached input at $0.6877 per 1M. Reasoning tokens are billed at the output rate.

Text, images and video in, text out, with up to 1,000,000 tokens in a single request and up to 131,072 tokens of output.

Send a POST request to https://api.aimlapi.com/v1/chat/completions with alibaba/qwen3.8-max-prime as the model id and your AI/ML API key in the Authorization header. The endpoint is OpenAI-compatible, so an existing chat client only needs the base URL swapped.

Start building with Qwen3.8 Max Prime

Get API Key
1000+ models, one API.