Nemotron Nano 12B V2 VL API

Nano 12B V2 VL is a 12-billion-parameter open multimodal reasoning model engineered for vision-language inference, processing text and multi-image inputs to generate coherent natural-language responses.
Output

How to use Nemotron Nano 12B V2 VL API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to .

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Nemotron Nano 12B V2 VL API Pricing

TypePrice
Input
Output

Nemotron Nano 12B V2 VL vs other models

ModelInputOutputContextBest for
$6.5 / 1M tokens
$39 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1.05M tokens
Reasoning + agents

Frequently asked questions

Yes, Nemotron Nano 12B V2 VL can stream responses as they are generated.

Yes, Nemotron Nano 12B V2 VL supports both function calling and structured outputs.

Nemotron Nano 12B V2 VL was built by NVIDIA.

The model is accessed via the v1/chat/completions endpoint.

Start building with Nemotron Nano 12B V2 VL

Get API Key
1000+ models, one API.