GLM-4.6 API

zhipu/glm-4.6
GLM-4.6 represents the cutting edge in large language models from Zhipu AI, balancing expansive context capabilities, efficient token use, and strong reasoning performance.
Context
200K tokens
Input
$0.78 / 1M tokens
Output
$2.86 / 1M tokens
Released
Oct 2, 2025

How to use GLM-4.6 API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to zhipu/glm-4.6.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "zhipu/glm-4.6",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "zhipu/glm-4.6",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"zhipu/glm-4.6","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

GLM-4.6 API Pricing

TypePrice
Input
$0.78 / 1M tokens
Output
$2.86 / 1M tokens
Cached input
$0.143 / 1M tokens

GLM-4.6 Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
LiveCodeBench
82.8%
Contamination-free competitive programming problemsSourceJuly 8, 2026
AIME 2025
93.9%
Competition mathematics (AIME), 2025SourceJuly 8, 2026
SWE-bench Verified
68%
Resolving verified real GitHub issuesSourceJuly 8, 2026
Intelligence
18.5
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Coding
45.8
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Math
86
Composite score across standardised mathematics evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026

GLM-4.6 vs other models

ModelInputOutputContextBest for
GLM-4.6
This page
$0.78 / 1M tokens
$2.86 / 1M tokens
200K tokens
Coding + agents
$6.5 / 1M tokens
$39 / 1M tokens
1M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$1.95 / 1M tokens
$11.7 / 1M tokens
1M tokens
Reasoning + agents

Frequently asked questions

GLM-4.6 has a 200,000 tokens context window and can return up to 128,000 tokens.

GLM-4.6 takes text as input and returns text.

Use zhipu/glm-4.6 as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

GLM-4.6 became available on October 2, 2025.

GLM-4.6 is priced at input $0.78 / 1M tokens, output $2.86 / 1M tokens, cached input $0.143 / 1M tokens.

Yes, GLM-4.6 can stream responses as they are generated.

GLM-4.6 was built by Zhipu AI.

Start building with GLM-4.6

Get API Key
1000+ models, one API.