import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "google/gemma-3-4b-it", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "google/gemma-3-4b-it", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"google/gemma-3-4b-it","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Intelligence | 4.8 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 22, 2026 |
| Coding | 2.7 | Composite score across standardised coding evaluations, measured independently by Artificial Analysis | Source | September 22, 2026 |
| Math | 12.7 | Composite score across standardised mathematics evaluations, measured independently by Artificial Analysis | Source | September 22, 2026 |
| MMLU | 74.5 | Multi-discipline knowledge across 57 school and professional subjects | Source | September 22, 2026 |
| MMLU-Pro | 45.3 | Multi-discipline knowledge + reasoning (harder MMLU) | Source | September 22, 2026 |
| MATH | 43.3 | Competition mathematics problems across five difficulty levels | Source | September 22, 2026 |
| GSM8K | 71.0 | Grade-school maths word problems requiring multi-step arithmetic | Source | September 22, 2026 |
| HumanEval | 45.7 | Python function synthesis from a docstring, scored by unit tests (pass@1) | Source | September 22, 2026 |
| DocVQA | 82.3 | Question answering over document images | Source | September 22, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Gemma 3 (4B) This page | Chat + assistants | |||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
Gemma 3 (4B) has a 131,000 tokens context window and can return up to 32,768 tokens.
Gemma 3 (4B) takes image, text as input and returns text.
Use google/gemma-3-4b-it as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.
Gemma 3 (4B) became available on September 9, 2025.
Gemma 3 (4B) is priced at input $0.06877 / 1M tokens, output $0.13754 / 1M tokens, cached input $0.055016 / 1M tokens.
Yes, Gemma 3 (4B) can stream responses as they are generated.
Yes, Gemma 3 (4B) accepts image input alongside text.
Gemma 3 (4B) was built by Google.
Yes, it is designed for on-device deployment and efficient inference.