NVIDIA Nemotron 3 Ultra is the largest model in the Nemotron 3 family: 550B parameters in total, 55B active per token, built as a hybrid mixture-of-experts (Mamba-2 and MoE layers with some attention layers), per NVIDIA's model card. InferAll serves it as nvidia/nemotron-3-ultra-550b-a55b at $0 input and $0 output on the open-model tier.
What we measured (2026-09-24, through InferAll)
| Test | Result |
|---|---|
Short reply, /v1/chat/completions, 3 calls |
200 each, 0.7s to 3.7s, answered as itself |
| Coding prompt (below), 80 output tokens | 2.6s |
Short reply, /v1/messages (Anthropic format) |
0.55s |
| Streaming, 17 chunks | first bytes 1.5s, complete 1.9s |
NVIDIA lists a context window of up to 1M tokens. We have not tested long contexts through InferAll, so treat that as NVIDIA's number, not ours. Like every model on NVIDIA's shared capacity, it can be briefly unavailable at busy times; InferAll then answers from another free model and says so in the response's model field, so check that field if the exact model matters.
Quick start (Python)
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_key_here", # get one at inferall.ai/keys
)
response = client.chat.completions.create(
model="nvidia/nemotron-3-ultra-550b-a55b",
messages=[{"role": "user", "content": "Write a Python function is_palindrome(s) that ignores case and non-alphanumeric characters. Code only."}],
max_tokens=600,
)
print(response.choices[0].message.content)
print(response.model) # the model that actually answered
What it returned in our test:
def is_palindrome(s: str) -> bool:
filtered = [c.lower() for c in s if c.isalnum()]
return filtered == filtered[::-1]
TypeScript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.inferall.ai/v1",
apiKey: process.env.INFERALL_API_KEY,
});
const response = await client.chat.completions.create({
model: "nvidia/nemotron-3-ultra-550b-a55b",
messages: [{ role: "user", content: "Explain this regex: ^\\d{3}-\\d{4}$" }],
max_tokens: 400,
});
console.log(response.choices[0].message.content);
Claude Code
export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=nvidia/nemotron-3-ultra-550b-a55b
claude
Claude Code may warn that the id is not in its model catalog; that is expected for any non-Anthropic model and does not stop it working.
Nemotron 3 Ultra or Nemotron 3 Super?
nvidia/nemotron-3-super-120b-a12b (120B total, 12B active) is our default free model and is lighter. Ultra is the larger model; in our tests on 2026-09-24 it was also the faster of the two to answer correctly. Both are $0. Run the same prompt on each and compare.
Get a key
Sign up at inferall.ai/keys. New accounts get 25 free calls on open models, no card needed. The full live list is at GET /ai/v1/models.
Model facts: NVIDIA model card, NVIDIA Nemotron 3 Ultra. License: OpenMDW-1.1.