DeepSeek V4.1 Flash API

deepseek/deepseek-v4.1-flash
DeepSeek V4.1 Flash is the current DeepSeek Flash build — a 1M-context chat and reasoning model with thinking mode enabled by default and image understanding.
Context
1M tokens
Input
$0.39 / 1M tokens
Output
$1.56 / 1M tokens
Released
Sep 10, 2026

How to use DeepSeek V4.1 Flash API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to deepseek/deepseek-v4.1-flash.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "deepseek/deepseek-v4.1-flash",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "deepseek/deepseek-v4.1-flash",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-v4.1-flash","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

DeepSeek V4.1 Flash API Pricing

TypePrice
Input
$0.39 / 1M tokens
Output
$1.56 / 1M tokens
Cached input
$0.0078 / 1M tokens

DeepSeek V4.1 Flash Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
39.5
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026

DeepSeek V4.1 Flash vs other models

ModelInputOutputContextBest for
$0.39 / 1M tokens
$1.56 / 1M tokens
1M tokens
Reasoning + agents
$1.17 / 1M tokens
$3.51 / 1M tokens
1M tokens
Reasoning + agents
$6.5 / 1M tokens
$39 / 1M tokens
1M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows

Frequently asked questions

DeepSeek V4.1 Flash is the current build in DeepSeek's Flash line: a chat and reasoning model with a 1M-token context window, thinking mode enabled by default and image understanding.

DeepSeek reports that V4.1 Flash surpasses V4 Pro on performance, cost and speed. It also accepts image input, which V4 Pro does not.

Yes. It takes both image and text input and returns text, so screenshots, diagrams and scanned pages can go into the same request as your prompt.

Yes. Thinking mode is enabled by default, so the model reasons through a problem before answering unless you tell it otherwise.

1,048,576 tokens, with up to 384,000 tokens of output.

Yes, it supports function calling and structured outputs, alongside streaming responses.

Work that needs a very large window: whole-repository code analysis, long-document review, and agent loops that accumulate a lot of context. Image input widens that to mixed text-and-visual material.

Yes. The model also answers to the alias deepseek-flash.

Send a request to the chat completions endpoint with the model id deepseek/deepseek-v4.1-flash. The request shape is the same as for any other chat model on AI/ML API.

Start building with DeepSeek V4.1 Flash

Get API Key
1000+ models, one API.