Whisper Base API

deepgram/whisper-base
Whisper: Multilingual speech recognition model, robust, versatile, open-source.
Output
$0.00007583 / sec
Released
Dec 30, 2025

How to use Whisper Base API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to deepgram/whisper-base.
import requests, time

headers = {"Authorization": "Bearer " + AIMLAPI_KEY}
job = requests.post(
    "https://api.aimlapi.com/v1/stt/create",
    headers=headers,
    json={
      "model": "deepgram/whisper-base",
      "url": "https://example.com/audio.mp3"
    },
).json()
gid = job["generation_id"]

while True:
    res = requests.get(f"https://api.aimlapi.com/v1/stt/{gid}", headers=headers).json()
    if res.get("status") in ("completed", "error", "failed"):
        break
    time.sleep(3)
print(res)
const headers = {
  Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
  "Content-Type": "application/json",
};
const job = await (await fetch("https://api.aimlapi.com/v1/stt/create", {
  method: "POST",
  headers,
  body: JSON.stringify({
    "model": "deepgram/whisper-base",
    "url": "https://example.com/audio.mp3"
  }),
})).json();

let res;
do {
  await new Promise((r) => setTimeout(r, 3000));
  res = await (await fetch(`https://api.aimlapi.com/v1/stt/${job.generation_id}`, { headers })).json();
} while (!["completed", "error", "failed"].includes(res.status));
console.log(res);
# submit the job — the response contains "generation_id"
curl -X POST https://api.aimlapi.com/v1/stt/create \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepgram/whisper-base","url":"https://example.com/audio.mp3"}'

# then poll for the result until it is ready
curl "https://api.aimlapi.com/v1/stt/{generation_id}" -H "Authorization: Bearer $AIMLAPI_KEY"

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Whisper Base API Pricing

TypePrice
Output
$0.00007583 / sec

Whisper Base Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
FLEURS WER
12.4
WER on multilingual speech (102 languages)SourceSeptember 21, 2026
LibriSpeech WER
196.8
Word Error Rate on read English speech (lower better)SourceSeptember 21, 2026

Whisper Base vs other models

ModelInputOutputContextBest for
Whisper Base
This page
$0.00007583 / sec
Transcription
$130 / 1M characters
$130 / 1M
Speech synthesis
$130 / 1M characters
$0.13 / 1M tokens
Speech synthesis
$78 / 1M characters
$0.078 / 1M tokens
Speech synthesis
$13 / 1M characters
$0.013 / 1M tokens
Speech synthesis

Frequently asked questions

Whisper Base takes audio as input and returns text.

Whisper Base became available on December 30, 2025.

Whisper Base is priced at output $0.00007583 / sec.

Whisper Base was built by Deepgram.

Use deepgram/whisper-base as the model id on AI/ML API.

Yes. Whisper Base is served through AI/ML API, so the same key and endpoint format used for other models applies.

Yes, Whisper Base is described as a multilingual speech recognition model.

Yes, Deepgram describes Whisper Base as an open-source speech recognition model.

Start building with Whisper Base

Get API Key
1000+ models, one API.