Models / Free
Nemotron 3.5 Lightning free API
NVIDIA's Nemotron 3.5 Lightning: 30B parameters in total, 3B active per token.
On InferAll it costs $0 per input and output token. A new trial key can call it without a card; paid accounts can call it too. Send this id as the model:
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="https://api.inferall.ai/v1",
api_key="ifu_your_key_here", # inferall.ai/keys
)
response = client.chat.completions.create(
model="nvidia/nemotron-3.5-lightning-30b-a3b",
messages=[{"role": "user", "content": "Explain what an API is in one sentence."}],
max_tokens=300,
)
print(response.choices[0].message.content)curl
curl https://api.inferall.ai/v1/chat/completions \
-H "Authorization: Bearer $INFERALL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "nvidia/nemotron-3.5-lightning-30b-a3b", "messages": [{"role": "user", "content": "Hello"}], "max_tokens": 200}'Claude Code and the Anthropic SDK
The same key works on the Anthropic-compatible /v1/messages endpoint:
export ANTHROPIC_BASE_URL=https://api.inferall.ai
export ANTHROPIC_API_KEY=ifu_your_key_here
export ANTHROPIC_MODEL=nvidia/nemotron-3.5-lightning-30b-a3b
claudeGood to know
- Free models run on NVIDIA's shared capacity. If this model is busy or slow, InferAll may answer with another free model rather than make you wait, and the response's
modelfield says which one answered. - The live list of models and prices is at /ai/v1/models.