Models / Free
Nemotron 3 Super free API
NVIDIA's Nemotron 3 Super: 120B parameters in total, 12B active per token (mixture of experts).
On InferAll it costs $0 per input and output token. A new trial key can call it without a card; paid accounts can call it too. Send this id as the model:
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_key_here", # inferall.ai/keys
)
response = client.chat.completions.create(
model="nvidia/nemotron-3-super-120b-a12b",
messages=[{"role": "user", "content": "Explain what an API is in one sentence."}],
max_tokens=300,
)
print(response.choices[0].message.content)curl
curl https://api.inferall.ai/v1/chat/completions \
-H "Authorization: Bearer $INFERALL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "nvidia/nemotron-3-super-120b-a12b", "messages": [{"role": "user", "content": "Hello"}], "max_tokens": 200}'Claude Code and the Anthropic SDK
The same key works on the Anthropic-compatible /v1/messages endpoint:
export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=nvidia/nemotron-3-super-120b-a12b
claudeGood to know
- Free models run on NVIDIA's shared capacity. If this model is busy or slow, InferAll may answer with another free model rather than make you wait, and the response's
modelfield says which one answered. - The live list of models and prices is at /ai/v1/models.