Models / Free
Nemotron 3 Ultra free API
NVIDIA's Nemotron 3 Ultra: 550B parameters in total, 55B active per token. The largest free model InferAll serves.
On InferAll it costs $0 per input and output token. A new trial key can call it without a card; paid accounts can call it too. Send this id as the model:
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_key_here", # inferall.ai/keys
)
response = client.chat.completions.create(
model="nvidia/nemotron-3-ultra-550b-a55b",
messages=[{"role": "user", "content": "Explain what an API is in one sentence."}],
max_tokens=300,
)
print(response.choices[0].message.content)curl
curl https://api.inferall.ai/v1/chat/completions \
-H "Authorization: Bearer $INFERALL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "nvidia/nemotron-3-ultra-550b-a55b", "messages": [{"role": "user", "content": "Hello"}], "max_tokens": 200}'Claude Code and the Anthropic SDK
The same key works on the Anthropic-compatible /v1/messages endpoint:
export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=nvidia/nemotron-3-ultra-550b-a55b
claudeGood to know
- Free models run on NVIDIA's shared capacity. If this model is busy or slow, InferAll may answer with another free model rather than make you wait, and the response's
modelfield says which one answered. - The live list of models and prices is at /ai/v1/models.