← Blog

NVIDIA Nemotron 3 Ultra 550B: free API, OpenAI-compatible

Call NVIDIA Nemotron 3 Ultra (550B total, 55B active) for free through InferAll's OpenAI- and Anthropic-compatible endpoints. Measured latency, working Python, TypeScript and Claude Code setup.

InferAll Team

3 min read
NVIDIA NIMNemotronNemotron 3 Ultrafree LLM APIOpenAI APIopen sourcedeveloper tools

NVIDIA Nemotron 3 Ultra is the largest model in the Nemotron 3 family: 550B parameters in total, 55B active per token, built as a hybrid mixture-of-experts (Mamba-2 and MoE layers with some attention layers), per NVIDIA's model card. InferAll serves it as nvidia/nemotron-3-ultra-550b-a55b at $0 input and $0 output on the open-model tier.

What we measured (2026-09-24, through InferAll)

Test Result
Short reply, /v1/chat/completions, 3 calls 200 each, 0.7s to 3.7s, answered as itself
Coding prompt (below), 80 output tokens 2.6s
Short reply, /v1/messages (Anthropic format) 0.55s
Streaming, 17 chunks first bytes 1.5s, complete 1.9s

NVIDIA lists a context window of up to 1M tokens. We have not tested long contexts through InferAll, so treat that as NVIDIA's number, not ours. Like every model on NVIDIA's shared capacity, it can be briefly unavailable at busy times; InferAll then answers from another free model and says so in the response's model field, so check that field if the exact model matters.


Quick start (Python)

from openai import OpenAI

client = OpenAI(
    base_url="https://api.inferall.ai/v1",
    api_key="ifu_your_key_here",  # get one at inferall.ai/keys
)

response = client.chat.completions.create(
    model="nvidia/nemotron-3-ultra-550b-a55b",
    messages=[{"role": "user", "content": "Write a Python function is_palindrome(s) that ignores case and non-alphanumeric characters. Code only."}],
    max_tokens=600,
)
print(response.choices[0].message.content)
print(response.model)  # the model that actually answered

What it returned in our test:

def is_palindrome(s: str) -> bool:
    filtered = [c.lower() for c in s if c.isalnum()]
    return filtered == filtered[::-1]

TypeScript

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.inferall.ai/v1",
  apiKey: process.env.INFERALL_API_KEY,
});

const response = await client.chat.completions.create({
  model: "nvidia/nemotron-3-ultra-550b-a55b",
  messages: [{ role: "user", content: "Explain this regex: ^\\d{3}-\\d{4}$" }],
  max_tokens: 400,
});
console.log(response.choices[0].message.content);

Claude Code

export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=nvidia/nemotron-3-ultra-550b-a55b
claude

Claude Code may warn that the id is not in its model catalog; that is expected for any non-Anthropic model and does not stop it working.


Nemotron 3 Ultra or Nemotron 3 Super?

nvidia/nemotron-3-super-120b-a12b (120B total, 12B active) is our default free model and is lighter. Ultra is the larger model; in our tests on 2026-09-24 it was also the faster of the two to answer correctly. Both are $0. Run the same prompt on each and compare.

Get a key

Sign up at inferall.ai/keys. New accounts get 25 free calls on open models, no card needed. The full live list is at GET /ai/v1/models.

Model facts: NVIDIA model card, NVIDIA Nemotron 3 Ultra. License: OpenMDW-1.1.

Try it with one key: create a free account and your first 25 calls on open models are free, no card needed.

Start building free