import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "fireworks/ember-1", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "fireworks/ember-1", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"fireworks/ember-1","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
Reasoning tokens are billed at the output rate — which is the point of the model: shorter traces mean fewer of them per answer.
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| SWE-bench Verified | 92.2% | Resolving verified real GitHub issues | Source | September 24, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Ember-1 This page | Reasoning with fewer thinking tokens | |||
| Long-context, multimodal & agentic workflows | ||||
| Frontier reasoning + agents | ||||
| Balanced coding + agents | ||||
| Reasoning + agents |
Ember-1 is a reasoning model from Fireworks Research, built on Moonshot AI's Kimi K3. It is tuned to produce shorter reasoning traces — to reach an answer while spending fewer thinking tokens on the way there.
On AI/ML API it costs $4.1262 per 1M input tokens and $20.631 per 1M output tokens, with cached input at $0.41262 per 1M. Reasoning tokens are billed at the output rate.
Kimi K3 is the base model Ember-1 is built on, at $3.9 / $19.5 per 1M tokens. Ember-1 lists slightly higher at $4.1262 / $20.631 but is tuned to use fewer reasoning tokens per answer, so which one is cheaper depends on how much your workload makes the model think. Kimi K3 also offers web search, which Ember-1 does not.
1,048,576 tokens in and up to 943,718 tokens out, inherited from Kimi K3 — enough for repository-scale code and very long replies in a single request.
Send a POST request to https://api.aimlapi.com/v1/chat/completions with fireworks/ember-1 as the model id and your AI/ML API key in the Authorization header. The endpoint is OpenAI-compatible, so an existing chat client only needs the base URL swapped.