Qwen3 VL Flash API

alibaba/qwen3-vl-flash
Qwen3 VL Flash blends fast multimodal vision-language processing with efficient memory usage, making it ideal for applications requiring cost-effective yet powerful visual reasoning.
Context
262K tokens
Input
$0.065 / 1M tokens
Output
$0.52 / 1M tokens
Released
Oct 19, 2025

How to use Qwen3 VL Flash API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to alibaba/qwen3-vl-flash.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "alibaba/qwen3-vl-flash",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "alibaba/qwen3-vl-flash",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"alibaba/qwen3-vl-flash","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Qwen3 VL Flash API Pricing

TypePrice
Input
$0.065 / 1M tokens
Output
$0.52 / 1M tokens
Cached input
$0.065 / 1M tokens

Qwen3 VL Flash vs other models

ModelInputOutputContextBest for
$0.065 / 1M tokens
$0.52 / 1M tokens
262K tokens
Reasoning + agents
$6.5 / 1M tokens
$39 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1.05M tokens
Reasoning + agents

Frequently asked questions

Qwen3 VL Flash has a 262,144 tokens context window and can return up to 32,768 tokens.

Qwen3 VL Flash takes image, text as input and returns text.

Use alibaba/qwen3-vl-flash as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Qwen3 VL Flash became available on October 19, 2025.

Qwen3 VL Flash is priced at input $0.065 / 1M tokens, output $0.52 / 1M tokens, cached input $0.065 / 1M tokens.

Yes, Qwen3 VL Flash can stream responses as they are generated.

Yes, Qwen3 VL Flash accepts image input alongside text.

Qwen3 VL Flash was built by Alibaba Cloud.

Start building with Qwen3 VL Flash

Get API Key
1000+ models, one API.