Gemma 3 (4B) API

google/gemma-3-4b-it
Multimodal AI model excelling in text and vision processing.​
Context
128K tokens
Input
$0.06877 / 1M tokens
Output
$0.13754 / 1M tokens
Released
Sep 9, 2025

How to use Gemma 3 (4B) API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to google/gemma-3-4b-it.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "google/gemma-3-4b-it",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "google/gemma-3-4b-it",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"google/gemma-3-4b-it","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Gemma 3 (4B) API Pricing

TypePrice
Input
$0.06877 / 1M tokens
Output
$0.13754 / 1M tokens
Cached input
$0.055016 / 1M tokens

Gemma 3 (4B) Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
4.8
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 22, 2026
Coding
2.7
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 22, 2026
Math
12.7
Composite score across standardised mathematics evaluations, measured independently by Artificial AnalysisSourceSeptember 22, 2026
MMLU
74.5
Multi-discipline knowledge across 57 school and professional subjectsSourceSeptember 22, 2026
MMLU-Pro
45.3
Multi-discipline knowledge + reasoning (harder MMLU)SourceSeptember 22, 2026
MATH
43.3
Competition mathematics problems across five difficulty levelsSourceSeptember 22, 2026
GSM8K
71.0
Grade-school maths word problems requiring multi-step arithmeticSourceSeptember 22, 2026
HumanEval
45.7
Python function synthesis from a docstring, scored by unit tests (pass@1)SourceSeptember 22, 2026
DocVQA
82.3
Question answering over document imagesSourceSeptember 22, 2026

Gemma 3 (4B) vs other models

ModelInputOutputContextBest for
Gemma 3 (4B)
This page
$0.06877 / 1M tokens
$0.13754 / 1M tokens
128K tokens
Chat + assistants
$6.5 / 1M tokens
$39 / 1M tokens
1M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$1.95 / 1M tokens
$11.7 / 1M tokens
1M tokens
Reasoning + agents

Frequently asked questions

Gemma 3 (4B) has a 131,000 tokens context window and can return up to 32,768 tokens.

Gemma 3 (4B) takes image, text as input and returns text.

Use google/gemma-3-4b-it as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Gemma 3 (4B) became available on September 9, 2025.

Gemma 3 (4B) is priced at input $0.06877 / 1M tokens, output $0.13754 / 1M tokens, cached input $0.055016 / 1M tokens.

Yes, Gemma 3 (4B) can stream responses as they are generated.

Yes, Gemma 3 (4B) accepts image input alongside text.

Gemma 3 (4B) was built by Google.

Yes, it is designed for on-device deployment and efficient inference.

Start building with Gemma 3 (4B)

Get API Key
1000+ models, one API.